vLLM repo
Every row links to its source: receipts, not summaries.
40Sightings
First seen
Last seen
- IQuest-Q1: How to Self-Host the 320B Agentic Coding Modelmindstudio.ai
- Building a High-Performance and Portable vLLM Linear Backend with Helionpytorch.org
- How to Deploy IQuest-Q1 with SGLang or vLLMmindstudio.ai
- How to Run Audio8 ASR Infinite Locally with vLLM or Dockermindstudio.ai
- How to Run OrcaSAQ2 27B on a 16GB GPUmindstudio.ai
- Watermarking in vLLMvllm.ai
- Inferact's TPU megakernel hits 709 tokens per second on Kimi K3runtimewire.com
- How to Run MiMo-V2.6-Flash-RL Locally with vLLM or SGLangmindstudio.ai
- Hardware-Agnostic Models in vLLMpytorch.org
- How Shopify built a continual learning loop with PyTorch and vLLMpytorch.org
- Paradigma releases Limite, a 1B model for high-throughput math reasoningruntimewire.com
- PD Serving of Qwen3.8-2.4Tvllm.ai
- Qwen-Image-2.1 on vLLM — Serve command for GB300 NVL4recipes.vllm.ai
- Scaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLMvllm.ai
- vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300vllm.ai
- Tiered KV Cache Offloading in vLLMvllm.ai
- GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLMvllm.ai
- vLLM x AgentX: Optimizing for Real-World Agentic Servingvllm.ai
- Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Pluginvllm.ai
- GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLMvllm.ai
- NVIDIA says new optimizations make local agents up to 1.9x fasterruntimewire.com
- NVIDIA Brings Simplified Local AI Support To NVIDIA GPUs Carrying 24+ GB VRAM While vLLM & & llama.cpp Optimizations Boost Compute By Up To 1.9xwccftech.com
- MiniMax H3 generates a 10.125-second audiovisual file in under nine secondsruntimewire.com
- How to Run GLM-5.3 Locally: Hardware, VRAM & Setup (2026)buildfastwithai.com
- MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3vllm.ai
- vLLM Sessions at PyTorch Conference North America 2026pytorch.org
- Exploring Speculative Decoding in vLLM on AMD GPUsvllm.ai
- IsoExec: Unified Execution to Eliminate Trainer-Inference Mismatch in SkyRLvllm.ai
- Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omnivllm.ai
- Adaptive Verification in vLLM: DSpark confidence-scheduled verificationvllm.ai