← AI Hub
vLLM
repo
Every row links to its source: receipts, not summaries.
12
Sightings
2026-06-25 20:34 UTC
First seen
2026-08-17 00:00 UTC
Last seen
2026-08-17 00:00 UTC
—
Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni
vllm.ai
2026-08-14 00:00 UTC
—
Adaptive Verification in vLLM: DSpark confidence-scheduled verification
vllm.ai
2026-08-12 00:00 UTC
—
Day 0 Support for Qwen3.8-2.4T-A95B on vLLM
vllm.ai
2026-08-03 00:00 UTC
—
Photon 2.0: Inference engine for Physical AI | Moondream
moondream.ai
2026-07-31 03:38 UTC
—
How to Run Kimi K3 Locally: Weights & Hardware (2026)
buildfastwithai.com
2026-07-23 00:47 UTC
—
Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis
newsletter.semianalysis.com
2026-07-08 23:50 UTC
—
Hugging Face Speeds Transformers Inference in vLLM
letsdatascience.com
2026-07-08 00:00 UTC
—
Native-speed vLLM transformers modeling backend
huggingface.co
2026-06-30 20:04 UTC
—
Article Compares Continuous and Static Batching in LLM Inference
letsdatascience.com
2026-06-28 04:58 UTC
—
Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference
marktechpost.com
2026-06-26 00:00 UTC
—
Run a vLLM Server on HF Jobs in One Command
huggingface.co
2026-06-25 20:34 UTC
—
Hugging Face Streamlines VLLM Deployment via HF Jobs in 2026
business20channel.tv