AiHub.
Entity timeline

vLLM repo

Every row links to its source: receipts, not summaries.

40Sightings
First seen
Last seen
  1. IQuest-Q1: How to Self-Host the 320B Agentic Coding Modelmindstudio.ai
  2. Building a High-Performance and Portable vLLM Linear Backend with Helionpytorch.org
  3. How to Deploy IQuest-Q1 with SGLang or vLLMmindstudio.ai
  4. How to Run Audio8 ASR Infinite Locally with vLLM or Dockermindstudio.ai
  5. How to Run OrcaSAQ2 27B on a 16GB GPUmindstudio.ai
  6. Watermarking in vLLMvllm.ai
  7. Inferact's TPU megakernel hits 709 tokens per second on Kimi K3runtimewire.com
  8. How to Run MiMo-V2.6-Flash-RL Locally with vLLM or SGLangmindstudio.ai
  9. Hardware-Agnostic Models in vLLMpytorch.org
  10. How Shopify built a continual learning loop with PyTorch and vLLMpytorch.org
  11. Paradigma releases Limite, a 1B model for high-throughput math reasoningruntimewire.com
  12. PD Serving of Qwen3.8-2.4Tvllm.ai
  13. Qwen-Image-2.1 on vLLM — Serve command for GB300 NVL4recipes.vllm.ai
  14. Scaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLMvllm.ai
  15. vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300vllm.ai
  16. Tiered KV Cache Offloading in vLLMvllm.ai
  17. GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLMvllm.ai
  18. vLLM x AgentX: Optimizing for Real-World Agentic Servingvllm.ai
  19. Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Pluginvllm.ai
  20. GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLMvllm.ai
  21. NVIDIA says new optimizations make local agents up to 1.9x fasterruntimewire.com
  22. NVIDIA Brings Simplified Local AI Support To NVIDIA GPUs Carrying 24+ GB VRAM While vLLM & & llama.cpp Optimizations Boost Compute By Up To 1.9xwccftech.com
  23. MiniMax H3 generates a 10.125-second audiovisual file in under nine secondsruntimewire.com
  24. How to Run GLM-5.3 Locally: Hardware, VRAM & Setup (2026)buildfastwithai.com
  25. MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3vllm.ai
  26. vLLM Sessions at PyTorch Conference North America 2026pytorch.org
  27. Exploring Speculative Decoding in vLLM on AMD GPUsvllm.ai
  28. IsoExec: Unified Execution to Eliminate Trainer-Inference Mismatch in SkyRLvllm.ai
  29. Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omnivllm.ai
  30. Adaptive Verification in vLLM: DSpark confidence-scheduled verificationvllm.ai