vLLM repo

Every row links to its source: receipts, not summaries.

12Sightings
First seen
Last seen
  1. Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni vllm.ai
  2. Adaptive Verification in vLLM: DSpark confidence-scheduled verification vllm.ai
  3. Day 0 Support for Qwen3.8-2.4T-A95B on vLLM vllm.ai
  4. Photon 2.0: Inference engine for Physical AI | Moondream moondream.ai
  5. How to Run Kimi K3 Locally: Weights & Hardware (2026) buildfastwithai.com
  6. Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis newsletter.semianalysis.com
  7. Hugging Face Speeds Transformers Inference in vLLM letsdatascience.com
  8. Native-speed vLLM transformers modeling backend huggingface.co
  9. Article Compares Continuous and Static Batching in LLM Inference letsdatascience.com
  10. Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference marktechpost.com
  11. Run a vLLM Server on HF Jobs in One Command huggingface.co
  12. Hugging Face Streamlines VLLM Deployment via HF Jobs in 2026 business20channel.tv