논문 리뷰: Kimi Delta Attention, 선형 어텐션이 전체 어텐션을 넘다 Kimi Linear - 채널별 게이팅과 하이브리드 구성으로 1M 컨텍스트를 감당하는 방법 Posted on July 21, 2026 논문 정보 [Read More] Tags: AI LLM Attention Paper-Review
Paper Review: Kimi Delta Attention — When Linear Attention Surpasses Full Attention Kimi Linear — handling a 1M context with channel-wise gating and a hybrid design Posted on July 21, 2026 Paper Info [Read More] Tags: AI LLM Attention Paper-Review
논문 리뷰: Attention Residuals, 잔차 연결을 학습된 어텐션으로 10년 된 고정 잔차를 계층 깊이에 대한 어텐션으로 대체하다 Posted on July 21, 2026 논문 정보 [Read More] Tags: AI LLM Transformer Paper-Review
Paper Review: Attention Residuals — Turning the Residual Connection into Learned Attention Replacing the decade-old fixed residual with attention over layer depth Posted on July 21, 2026 Paper Info [Read More] Tags: AI LLM Transformer Paper-Review
논문 리뷰: TaiChi, Prefill-Decode 통합으로 Goodput을 최적화하다 Aggregation이냐 Disaggregation이냐 - 둘을 합쳐 균형 SLO에서 이기는 LLM 서빙 시스템 Posted on July 20, 2026 논문 정보 [Read More] Tags: AI LLM Serving GPU Paper-Review