LIBRISTO
LIBROAMANTO
obligatoriu
Faceți parte dintr-o comunitate de iubitori de cărți din întreaga lume și beneficiați de o mulțime de avantaje Creați-vă un cont gratuit
0
Transport gratuit la punctele de livrare Pick Up peste 349.00 lei
Packeta 15.00 lei Serviciul de curierat Cargus 28.00 lei Easybox 20.00 lei FAN Courier 20.00 lei Punct FAN 16.00 lei Punct DPD 17.00 lei Curier DPD 25.00 lei

Livrare gratuită pentru comenzile peste 349,00 lei.

Inside LLM Inference

From Silicon to Served Token - A Measured Guide to How Language Models Run on GPUs

Limba englezăengleză
Carte Carte broșată
Carte Inside LLM Inference The AI Singularity
Codul Libristo: 53249500
Editura Independently published, iulie 2026
How does a language model actually become served tokens on a GPU - and why is it so much slower than... Descrierea completă
? points 57 b Curând Curând Nou Nou
123.81 lei
Așteptăm intrarea în stoc Ediția 20. 07. 2026

Până la 30 de zile pentru returnare

How does a language model actually become served tokens on a GPU - and why is it so much slower than the hardware's headline TFLOPS promise?

Inside LLM Inference follows a single request through the entire stack: from a PyTorch call down through the CUDA runtime, the streaming multiprocessor, the warp, and the memory system, to the arithmetic cores that finally do the work - and back up to the throughput and cost you pay for.

Every claim is measured, not asserted. A production model runs on 2026's newest silicon - NVIDIA's Blackwell generation - across four inference engines (vLLM, SGLang, TensorRT-LLM, and Hugging Face Transformers), profiled to individual GPU kernels with hardware counters and cross-checked against the physics that predicts them.

You start from zero - Chapter 0 defines every term - and finish able to read a live server:

  • classify a kernel as memory- or compute-bound, and prove it with a clock-sensitivity test
  • size a KV-cache pool and the batch it admits from a model's own dimensions
  • read sm_active and sm_occupancy off a running GPU and know what they mean
  • choose an inference engine on evidence, not vibes

The closing chapters reach the 2026 state of the art: speculative decoding and self-drafting (EAGLE-3, multi-token prediction), the compressed KV cache of multi-head latent attention, prefill/decode disaggregation, mixture-of-experts routing, and the Blackwell-era FlashAttention-4 kernel.

One thesis runs throughout: serving language models is fundamentally a memory-bandwidth problem, and batch size is the price of admission to compute you have already paid for.

Whether you deploy models, optimize inference, or simply want to understand what your GPU is really doing between tokens, this book turns the black box into something you can measure, predict, and tune.

Actriță & Poliglotă
EWA KASP pentru
Redă videoclipul
Ewa Kasp
Libristo are cea mai mare selecție de literatură în limbi străine. De aceea îmi cumpăr cărțile de aici.

Informații despre carte

Titlu complet Inside LLM Inference
Limba engleză
Legare Carte - Carte broșată
Data publicării 2026
Număr pagini 180
EAN 9798187180769
Codul Libristo 53249500
Greutatea 324
Dimensiuni 178 x 254 x 10
Dăruiește această carte chiar astăzi
Este foarte ușor
1 Adaugă cartea în coș și selectează Livrează ca un cadou 2 Îți vom trimite un voucher în schimb 3 Cartea va ajunge direct la adresa destinatarului

Logare

Conectare la contul de utilizator Încă nu ai un cont Libristo? Crează acum!

 
obligatoriu
obligatoriu

Nu ai un cont? Beneficii cu contul Libristo!

Datorită contului Libristo, vei avea totul sub control.

Creare cont Libristo
Consilier de cărți Libroamiko
Bună ziua, sunt Libroamiko, vă pot ajuta?