K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

AI Runtimes Advance with Apple Silicon Support, Agent Sandbox Streamlining

Published , 04:44 Bogota (UTC-5) · 15 sourced items, 3 new since the previous edition · Read the foundations review · RSS

Today in 5 points

Phone and edge AI

datasette-auth-github 1.0NEW

Simon Willison · · Phone and edge AI

datasette-auth-github 1.0 has been released, addressing an issue where authenticated sessions expired prematurely due to cookies lacking a Max-Age parameter, particularly affecting Mobile Safari.

Why it matters: This improves session stability for users interacting with Datasette demo sites or similar agent sandboxes that use GitHub authentication.

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

NVIDIA Technical Blog · · Phone and edge AI

TensorRT Edge-LLM completed the MLPerf Edge Agentic Benchmark 6.4 times faster on Jetson AGX Thor, indicating a trend of AI agents moving from cloud data centers to edge devices.

Why it matters: This highlights advancements in running AI agents efficiently on edge hardware, which is relevant for users considering deploying AI on devices like the Jetson AGX Thor.

v0.17.1

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.1 includes a bug fix addressing a tool call integer type issue.

Why it matters: This ensures correct tool call functionality within LiteRT-LM, which is important for local AI applications that integrate with external tools.

Claude Cowork and chat are now one Claude

Simon Willison · · Phone and edge AI

Claude Cowork and chat functionalities are merging into a single Claude experience, rolling out to Pro and Max plans on web, desktop, and mobile. This suggests Claude is evolving into a general agent.

Why it matters: This simplifies the user experience for Claude users across platforms, including mobile, and indicates a trend towards unified agent capabilities.

Runtimes and quantization

b11060NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp release b11060 includes fixes for Mamba, specifically making time-step projection input contiguous and skipping contiguous copy after normalization. It supports a wide range of platforms, including macOS Apple Silicon, Linux with various accelerators

Why it matters: This improves Mamba model performance and efficiency within llama.cpp, benefiting users running AI locally on diverse hardware, including Apple Silicon.

b11058NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp release b11058 provides a fix for Gemma4 required tool grammar.

Why it matters: This ensures correct functionality for Gemma4 models when using tool grammar within llama.cpp, which is relevant for local AI development.

v0.34.3

Ollama releases · · Runtimes and quantization

Ollama v0.34.3 adds support for Nemotron H vision models on Apple Silicon with MLX. The API now exposes model thinking controls and their defaults. The macOS app no longer reopens closed windows, and model pulls from HuggingFace are fixed.

Why it matters: This significantly enhances local AI capabilities on Apple Silicon by supporting new vision models and improves user experience with the Ollama macOS app.

v0.30.0rc2

vLLM releases · · Runtimes and quantization

vLLM v0.30.0rc2 includes a bugfix to prevent receive reports for notification-only requests.

Why it matters: This is a technical bugfix that contributes to the stability and efficiency of the vLLM inference runtime, benefiting users running large language models locally.

v0.34.3-rc0

Ollama releases · · Runtimes and quantization

Ollama's API now exposes model thinking levels and their default settings.

Why it matters: This provides more control and transparency for users interacting with Ollama models via its API, allowing for better configuration of local AI inference.

Where the Industry Is Investing: A Look at MLPerf Inference v6.1

MLCommons · · Runtimes and quantization

An analysis of MLPerf Inference v6.1 highlights a record number of submitters and systems, noting the introduction of agentic and end-to-end benchmarks.

Why it matters: This indicates industry trends in AI inference, including a focus on agentic workloads, which is relevant for understanding the evolving landscape of local AI.

v0.30.0rc1: [Bugfix] Isolate supplemental FlashInfer BF16 autotuning (#57285)

vLLM releases · · Runtimes and quantization

vLLM v0.30.0rc1 includes a bugfix to isolate supplemental FlashInfer BF16 autotuning.

Why it matters: This is a technical bugfix that contributes to the stability and efficiency of the vLLM inference runtime, benefiting users running large language models locally.

Agent sandboxes (E2B and peers)

e2b@2.51.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK e2b@2.51.0 removes SDK-side defaults from API request payloads, ensuring API defaults apply. Sandbox create, fork, and connect operations no longer preset various options. Create and connect now use v2 API endpoints, which default to a 5-minute timeout

Why it matters: This streamlines sandbox configuration by relying on API defaults, enhances security with v2 API endpoints, and simplifies the SDK for users developing agent sandboxes.

How Lark Uses E2B to Safely Test Apps with Customer Data

E2B Blog · · Agent sandboxes (E2B and peers)

Lark utilizes E2B sandboxes for safely testing customer applications within Docker-based development environments.

Why it matters: This demonstrates a practical application of E2B sandboxes for secure and isolated testing, which is relevant for users building and deploying agent sandboxes.

e2b@2.50.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK e2b@2.50.0 removed V1 template build operations and schemas from its generated API clients, as the API no longer serves them. Template builds are now handled through the Template SDK.

Why it matters: This indicates an API modernization for E2B, requiring users of the SDK to adapt to the new Template SDK for template builds, which is important for agent sandbox development.

Open models for local use

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

NVIDIA Technical Blog · · Open models for local use

Alibaba has released Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model with 2.4 trillion parameters and configurable reasoning, designed for serving on NVIDIA GB300 NVL72.

Why it matters: This introduces a large, open-weight model with advanced capabilities, relevant for users interested in deploying powerful models.

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive