K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF
Local AI Runtimes Advance with MLX Defaults, NPU Support, and Multi-GPU Inference
Published , 04:44 Bogota (UTC-5) · 9 sourced items, 4 new since the previous edition · Read the foundations review · RSS
Today in 5 points
Ollama now automatically runs models on MLX on Apple Silicon devices by default for supported architectures. This release also brings faster Qwen 3.8 prompt processing and improved Gemma 4 image resolution on Apple Silicon. [3][4]
llama.cpp has expanded its runtime capabilities, adding Hexagon NPU support for tiled Q4_0 and Q8_0 GET_ROWS, with improvements to DMA pipeline and kernel selection. It also supports Nemotron 3 Puzzle state size 96 for ssm scan in CUDA. [1][2]
vLLM v0.30.0 introduces several new models and features a persistent per-GPU weight-cache daemon for faster engine restarts. Additionally, vLLM has added MI355 dense NVFP4 and MoRI kernel mirrors for ROCm. [6][7]
Hugging Face Transformers now supports running llama.cpp quantized models. Separately, NVIDIA TensorRT multi-device inference is a new capability for serving models across multiple GPUs in NVIDIA Dynamo-Triton. [8][9]
K/20X research paper
Frontier Models, September 2026: ASTRA, Fable, Jev and the Chinese FrontierK/20X LABS paper · 2026-09-21 GPT-6 Astra and Claude Fable 5.1 tie at 53 on the AA Intelligence Index and split the specialised benchmarks; Qwen3.8 Max, GLM-5.3 and Kimi K3 trail by 8 to 9 points at about a quarter of the cost; TypeSafe's Jev returns typed decisions in under 500 ms at $0.042/M and unbundles classification work from frontier LLMs.
Apple Machine Learning Research · · Phone and edge AI
System-wide Dictation on Apple devices operates entirely on-device, using an encoder to map audio waveforms to the language model's representation.
Why it matters: This indicates Apple's focus on on-device AI for core system functions, which can inform local AI development strategies for privacy and efficiency.
llama.cpp now supports Hexagon for tiled Q4_0 and Q8_0 GET_ROWS, with improvements to DMA pipeline, kernel selection logic, and reenabled vectorizer for hex-build.
Why it matters: This expands llama.cpp's capabilities on Qualcomm Hexagon NPUs, potentially improving performance for specific quantized models on supported devices.
llama.cpp adds support for Nemotron 3 Puzzle state size 96 for ssm scan in CUDA and lists broad platform support including Snapdragon's Hexagon NPU on Linux and Android.
Why it matters: This extends llama.cpp's model compatibility and highlights its wide platform support, including mobile NPUs, for local inference.
Ollama improved structured outputs, fixed "model not found" errors and macOS app unresponsiveness, and updated core components. Qwen 3.8 prompt processing is faster on Apple Silicon, and Gemma 4 on Apple Silicon handles image resolution better.
Why it matters: These updates improve the reliability, performance, and user experience of Ollama for local AI inference on Apple Silicon devices.
vLLM has added MI355 dense NVFP4 and MoRI kernel mirrors for ROCm.
Why it matters: This indicates expanded support and optimization for AMD's ROCm platform, potentially benefiting users with AMD GPUs for local inference.
vLLM v0.30.0 introduces several new models, including DeepSeek-V4.1-Flash and DeepGEMM Mega-mHC, and features a persistent per-GPU weight-cache daemon for faster engine restarts.
Why it matters: This release expands the range of models available through vLLM and offers a feature to speed up engine restarts for multi-GPU setups.
Hugging Face Transformers now supports running llama.cpp quantized models.
Why it matters: This integration allows users to leverage llama.cpp's efficient quantized models directly within the Transformers library for local inference.
NVIDIA Technical Blog · · Runtimes and quantization
NVIDIA TensorRT multi-device inference is a new capability in NVIDIA Dynamo-Triton for serving models across multiple GPUs.
Why it matters: This feature addresses the demands of large generative AI models by enabling efficient inference across multiple NVIDIA GPUs.
Method
Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.