[모델] 4 layer / 4 head / 128 dim / context 128 -> 0.83M 파라미터, 어휘 65자
  iter    0 | train 4.334 | val 4.323 |    3.4s
  iter  250 | train 2.488 | val 2.489 |   62.9s
  iter  500 | train 2.352 | val 2.366 |  122.1s
  iter  750 | train 2.204 | val 2.209 |  181.3s
  iter 1000 | train 2.058 | val 2.101 |  241.2s
  iter 1250 | train 1.950 | val 2.030 |  302.8s
  iter 1500 | train 1.869 | val 1.971 |  360.8s
  iter 1750 | train 1.817 | val 1.937 |  418.2s
  iter 2000 | train 1.757 | val 1.880 |  477.2s
[학습] 2000 iter, 477s (CPU)

[생성 샘플]
--------------------------------------------------
Lole hap essend there I bese
And tegring pachined his grease, spiting's ruld
eale gids fitt reas in afs the fill are capur
As he him, and is upon to this no, chande in
till Love for to sonern.

For RIVt:
Ill see they sency seeact and a will sheen.

BERDIZUS:
Neathed crin'd of be and their and us dame,
And and that and'd and strosed many alt
Of distised of thes thine your-have heir is the proven
Th
--------------------------------------------------
[생성 속도] KV cache 미사용:  302.9 tok/s
[생성 속도] KV cache 사용:  543.0 tok/s
  -> KV cache로 1.79배 가속 (모델이 클수록 격차 커짐)

[KV cache 메모리] 시퀀스 120토큰, batch 1 기준 실측 480.0 KiB (= 2 x 4 layer x 128 dim x 120 x 4B)
  같은 식을 GPT-3 175B급(96 layer, 12288 dim, FP16)에 적용하면:
    컨텍스트  2,048 토큰 -> 사용자 1명당 KV cache    9.0 GiB
    컨텍스트 32,768 토큰 -> 사용자 1명당 KV cache  144.0 GiB
  -> 긴 컨텍스트 x 동시 사용자 수만큼 선형 증가. 추론 서버 메모리 용량을 압박하는 실체.