-
지식 베이스 구축: Cerebras 사례와 실전 목업
데이터를 옮기지 않고 연결하는 접근, 그리고 프롬프트로 테스트한 KB 파이프라인 제안
Posted on July 28, 2026
-
Building a Knowledge Base: Lessons from Cerebras and a Tested Mockup
Connect data where it lives, plus a prompt-tested KB pipeline proposal
Posted on July 28, 2026
-
RISC-V CPU #4: 파이프라인 심화, 해저드와 분기 예측
겹쳐 흐르는 명령어들의 충돌을 하드웨어가 어떻게 감지하고 해결하는가
Posted on July 22, 2026
-
RISC-V CPU #4: The Pipeline in Depth — Hazards and Branch Prediction
How hardware detects and resolves the conflicts between overlapping instructions
Posted on July 22, 2026
-
RISC-V CPU #3: 파이프라인의 거시 구조
명령어 하나가 거치는 다섯 단계와, 각 단계를 담당하는 하드웨어 유닛
Posted on July 21, 2026
-
RISC-V CPU #3: The Macro Structure of the Pipeline
The five stages one instruction passes through, and the hardware unit behind each
Posted on July 21, 2026
-
논문 리뷰: LatentMoE, FLOP·파라미터당 정확도를 끌어올리는 MoE
잠재공간에서 라우팅하는 MoE - 그리고 Kimi K3의 Stable LatentMoE
Posted on July 21, 2026
-
Paper Review: LatentMoE — Raising Accuracy per FLOP and Parameter in MoE
Routing an MoE in a latent space — and Kimi K3's Stable LatentMoE
Posted on July 21, 2026
-
Kimi K3 분석: 2.8조 파라미터 오픈웨이트 MoE, 어디까지 왔나
Moonshot AI의 최대 공개 모델 - 구조, 신기법, 운용 인프라, 성능을 다각도로
Posted on July 21, 2026
-
Analyzing Kimi K3: A 2.8-Trillion-Parameter Open-Weight MoE, How Far Has It Come
Moonshot AI's largest open model — architecture, new techniques, serving infrastructure, and performance
Posted on July 21, 2026
-
논문 리뷰: Kimi Delta Attention, 선형 어텐션이 전체 어텐션을 넘다
Kimi Linear - 채널별 게이팅과 하이브리드 구성으로 1M 컨텍스트를 감당하는 방법
Posted on July 21, 2026
-
Paper Review: Kimi Delta Attention — When Linear Attention Surpasses Full Attention
Kimi Linear — handling a 1M context with channel-wise gating and a hybrid design
Posted on July 21, 2026
-
논문 리뷰: Attention Residuals, 잔차 연결을 학습된 어텐션으로
10년 된 고정 잔차를 계층 깊이에 대한 어텐션으로 대체하다
Posted on July 21, 2026
-
Paper Review: Attention Residuals — Turning the Residual Connection into Learned Attention
Replacing the decade-old fixed residual with attention over layer depth
Posted on July 21, 2026
-
논문 리뷰: TaiChi, Prefill-Decode 통합으로 Goodput을 최적화하다
Aggregation이냐 Disaggregation이냐 - 둘을 합쳐 균형 SLO에서 이기는 LLM 서빙 시스템
Posted on July 20, 2026
-
Paper Review: TaiChi — Unifying Prefill-Decode for Goodput-Optimized LLM Serving
Aggregation or disaggregation? Combine both to win under balanced SLOs
Posted on July 20, 2026
-
RISC-V CPU #2: Instruction이란
명령어는 32비트 숫자다: 인코딩, 여섯 가지 포맷, 그리고 즉시값의 수수께끼
Posted on July 20, 2026
-
RISC-V CPU #2: What Is an Instruction
An instruction is a 32-bit number: encoding, six formats, and the puzzle of the immediate
Posted on July 20, 2026
-
RISC-V CPU #1: RISC-V란 무엇인가
ISA라는 계약, RISC 철학, 그리고 오픈 ISA로서의 RISC-V
Posted on July 19, 2026
-
RISC-V CPU #1: What Is RISC-V?
The ISA as a contract, the RISC philosophy, and RISC-V as an open ISA
Posted on July 19, 2026
-
AMD GPU 아키텍처 #7: CDNA 3
MI300X - 3D 적층, 통합 논리 GPU, 192GB HBM3, ROCm 6로 AI 추론 시장에 진입하다
Posted on July 16, 2026
-
AMD GPU Architecture #7: CDNA 3
MI300X — 3D stacking, a unified logical GPU, 192GB HBM3, and ROCm 6 entering the AI inference market
Posted on July 16, 2026
-
AMD GPU 아키텍처 #6: CDNA 2
AMD 최초 GPU MCM, FP64 Matrix Core, 그리고 세계 최초 엑사스케일 Frontier
Posted on July 15, 2026
-
AMD GPU Architecture #6: CDNA 2
AMD's first GPU MCM, the FP64 Matrix Core, and the world's first exascale system, Frontier
Posted on July 15, 2026
-
AMD GPU 아키텍처 #5: CDNA 1
GCN 컴퓨트 계보의 계승, Matrix Core의 등장, 그리고 NVIDIA A100과의 대결
Posted on July 15, 2026
-
AMD GPU Architecture #5: CDNA 1
Inheriting GCN's compute lineage, the arrival of Matrix Core, and a head-to-head with NVIDIA A100
Posted on July 15, 2026
-
AMD GPU 아키텍처 #4: RDNA 4
모놀리식 복귀, 3세대 Ray Accelerator, FSR 4의 ML 전환, 그리고 Blackwell과의 대결
Posted on July 13, 2026
-
AMD GPU Architecture #4: RDNA 4
Return to monolithic, 3rd-gen Ray Accelerators, FSR 4's ML shift, and a head-to-head with Blackwell
Posted on July 13, 2026
-
AMD GPU 아키텍처 #3: RDNA 3
칩렛으로 분리된 다이, Dual-Issue 셰이더, 그리고 NVIDIA Ada Lovelace와의 대결
Posted on July 13, 2026
-
AMD GPU Architecture #3: RDNA 3
The chiplet die split, dual-issue shaders, and a head-to-head with NVIDIA Ada Lovelace
Posted on July 13, 2026
-
AMD GPU 아키텍처 #2: RDNA 2
Ray Accelerator, Infinity Cache, Big Navi - 그리고 NVIDIA Ampere와의 정면 대결
Posted on July 13, 2026
-
AMD GPU Architecture #2: RDNA 2
Ray Accelerator, Infinity Cache, Big Navi — head-to-head with NVIDIA Ampere
Posted on July 13, 2026
-
Vera CPU / Rubin GPU: GTC 2025 공개 내용 정리
공식 발표 한도 내에서 정리한 NVIDIA 차세대 AI 플랫폼
Posted on July 9, 2026
-
Vera CPU / Rubin GPU: What NVIDIA Announced at GTC 2025
A summary of publicly confirmed details for NVIDIA's next-generation AI platform
Posted on July 9, 2026
-
AMD GPU 아키텍처 계보: RDNA와 CDNA
소비자용 RDNA와 데이터센터용 CDNA, 출시 순서로 정리
Posted on July 9, 2026
-
AMD GPU Architecture Lineage: RDNA and CDNA
Consumer RDNA and datacenter CDNA — every generation in chronological order
Posted on July 9, 2026
-
AMD GPU 아키텍처 #1: RDNA 1
Wave32·WGP·캐시 계층 개편, 그리고 NVIDIA Turing과의 구조적 비교
Posted on July 9, 2026
-
AMD GPU Architecture #1: RDNA 1
Wave32, WGP, cache redesign — and a structural comparison with NVIDIA Turing
Posted on July 9, 2026
-
GPU 아키텍처 #9: Blackwell - FP4 Tensor Core와 멀티다이 설계
MXFP8 미세조정, NV-HBI 2-다이 MCM, DLSS 4 다중 프레임 생성
Posted on July 8, 2026
-
GPU Architecture #9: Blackwell — FP4 Tensor Cores and Multi-Die Design
MXFP8 Microscaling, NV-HBI Two-Die MCM, and DLSS 4 Multi-Frame Generation
Posted on July 8, 2026
-
CXL이란 무엇인가 - Compute Express Link 개요
PCIe 위의 캐시 일관성 인터커넥트, 메모리 확장, AI/HPC 활용
Posted on July 8, 2026
-
What is CXL? — Compute Express Link Overview
Cache-coherent interconnect over PCIe, memory expansion, and AI/HPC use cases
Posted on July 8, 2026
-
Grace CPU 아키텍처: NVIDIA의 첫 데이터센터 CPU
Neoverse V2 72코어, LPDDR5X, NVLink-C2C로 구현한 CPU-GPU 통합
Posted on July 8, 2026
-
Grace CPU Architecture: NVIDIA's First Datacenter CPU
72 Neoverse V2 cores, LPDDR5X, and NVLink-C2C for unified CPU-GPU compute
Posted on July 8, 2026
-
InfiniBand와 RoCE: GPU 클러스터 고속 네트워크의 기반
RDMA 원리, Verbs API, Lossless Ethernet, DCQCN, GPUDirect RDMA
Posted on July 7, 2026
-
InfiniBand and RoCE — The Network Foundation of GPU Clusters
RDMA mechanics, Verbs API, lossless Ethernet, DCQCN, GPUDirect RDMA
Posted on July 7, 2026
-
GPU 아키텍처 #8: Ada Lovelace - 3세대 RT Core와 96MB L2
Opacity Micromap, Shader Execution Reordering, DLSS 3 Frame Generation
Posted on July 7, 2026
-
GPU Architecture #8: Ada Lovelace — 3rd-Gen RT Cores and 96MB L2
Opacity Micromap Engine, Shader Execution Reordering, and DLSS 3 Frame Generation
Posted on July 7, 2026
-
GPU 아키텍처 #7: Hopper - Transformer Engine과 FP8
4세대 Tensor Core의 FP8 가속, Transformer Engine, Thread Block Cluster, GH200 Grace Hopper
Posted on July 7, 2026
-
GPU Architecture #7: Hopper — Transformer Engine and FP8
4th-gen Tensor Cores with FP8, Transformer Engine, Thread Block Clusters, and GH200 Grace Hopper
Posted on July 7, 2026
-
GPU 아키텍처 #6: Ampere - Sparsity 가속과 MIG
3세대 Tensor Core의 2:4 구조적 희소성, TF32/BF16, A100 MIG, cp.async 파이프라이닝
Posted on July 7, 2026
-
GPU Architecture #6: Ampere — Sparsity Acceleration and MIG
3rd-gen Tensor Cores with 2:4 structured sparsity, TF32/BF16, A100 MIG, and cp.async pipelining
Posted on July 7, 2026
-
GPU 아키텍처 #5: Turing - RT Core와 2세대 Tensor Core
실시간 레이트레이싱 전용 하드웨어, INT8/INT4 Tensor Core, 그리고 소비자 GPU의 AI 가속
Posted on July 7, 2026
-
GPU Architecture #5: Turing — RT Cores and 2nd-Gen Tensor Cores
Hardware ray tracing acceleration, INT8/INT4 Tensor Cores, and AI inference on consumer GPUs
Posted on July 7, 2026
-
GPU 아키텍처 #4: Volta - Tensor Core와 독립 스레드 스케줄링
GV100 SM 재설계, FP32/INT32 분리 실행, 독립 스레드 스케줄링, WMMA API
Posted on July 7, 2026
-
GPU Architecture #4: Volta — Tensor Cores and Independent Thread Scheduling
GV100 SM redesign, FP32/INT32 dual pipelines, per-thread PC, and WMMA API
Posted on July 7, 2026
-
GPU 아키텍처 #3: Pascal - 16nm, HBM2, NVLink
FinFET 공정 전환, GP100 SM 재설계, Unified Memory의 도약
Posted on July 7, 2026
-
GPU Architecture #3: Pascal — 16nm, HBM2, NVLink
FinFET process transition, GP100 SM redesign, and Unified Memory's leap forward
Posted on July 7, 2026
-
GPU 아키텍처 #2: Kepler와 Maxwell - 효율성의 탐구
192코어 SMX의 실험, Quadrant 설계로의 수렴, 그리고 스케줄링의 진화
Posted on July 6, 2026
-
GPU Architecture #2: Kepler and Maxwell — The Pursuit of Efficiency
The 192-core SMX experiment, convergence to Quadrant design, and scheduling evolution
Posted on July 6, 2026
-
LLM 서빙 프레임워크 #1: vLLM 심층 분석
다중 프로세스 아키텍처, Scheduler, KVCacheManager, 병렬화 전략 (V1 기준)
Posted on June 22, 2026
-
LLM Serving Frameworks #1: Deep Dive into vLLM
Multi-process architecture, Scheduler, KVCacheManager, and parallelism strategies (V1)
Posted on June 22, 2026
-
LLM 서빙 프레임워크 개요: vLLM, SGLang, TensorRT-LLM
세 프레임워크의 설계 철학과 핵심 기술 비교
Posted on June 19, 2026
-
LLM Serving Frameworks Overview: vLLM, SGLang, TensorRT-LLM
Design philosophies and core technologies compared
Posted on June 19, 2026
-
GPU 아키텍처 #1: GPU의 출발과 SIMT의 탄생
Tesla, Fermi 아키텍처와 CUDA의 시작
Posted on April 12, 2026
-
GPU Architecture Series #1: The Origin of GPU and the Birth of SIMT
Tesla, Fermi Architecture and the Beginning of CUDA
Posted on April 12, 2026
No posts found.