K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

Local AI Runtimes Advance, Agent Sandboxes Enhance Security, and Edge AI Influences Mobile Dev

Published , 04:44 Bogota (UTC-5) · 29 sourced items, 25 new since the previous edition · Read the foundations review · RSS

Today in 5 points

Laptop AI (MacBook, MLX)

Desk-side boxes (Mac Studio, DGX Spark, OEM)

SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PCNEW

arXiv · · Desk-side boxes (Mac Studio, DGX Spark, OEM)

SSD-LLaMA is an SSD-native local MoE inference system enabling trillion-parameter models on consumer PCs by optimizing SSD I/O, managing a three-tier storage hierarchy, and balancing CPU-GPU execution.

Why it matters: Makes very large MoE models accessible for local inference on consumer-grade hardware by leveraging SSDs for capacity.

Phone and edge AI

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI AgentsNEW

arXiv · · Phone and edge AI

EvoSkill-GUI is a training-free framework enabling GUI agents to revise their skills from execution feedback at deployment time, without additional training.

Why it matters: Improves the adaptability and reliability of GUI agents on local devices by allowing skill evolution without retraining.

vidax: A Unified JAX Framework for Video Generative Models on Accelerator MeshesNEW

arXiv · · Phone and edge AI

vidax is an open-source JAX/Flax inference engine and PyTorch-to-JAX weight translator for video generative models, supporting Cloud TPU pods and integrating TPU flash-attention kernels.

Why it matters: Provides a production-ready inference path for video generative models on Cloud TPUs, expanding options for phone AI development.

Libra: Efficient Resource Management for Agentic RL Post-TrainingNEW

arXiv · · Phone and edge AI

Libra is an adaptive runtime for agentic RL post-training, designed to manage long-tailed and non-stationary workloads by balancing execution across rollout and training stages.

Why it matters: Improves resource management for agentic RL post-training, which can be critical for optimizing phone AI development.

v0.17.1NEW

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.1 has been released, including a bug fix for a tool call integer type issue.

Why it matters: Indicates ongoing development and bug fixes for a runtime focused on efficient, potentially mobile-oriented, LLM inference.

Claude Cowork and chat are now one ClaudeNEW

Simon Willison · · Phone and edge AI

Claude Cowork and chat are merging into a single Claude experience, which is described as becoming a general agent. This is rolling out to Pro and Max plans.

Why it matters: Reflects a trend towards unified, general-purpose AI agents, impacting how users interact with AI on mobile and desktop.

Native is now the future of mobile at Shopify

Simon Willison · · Phone and edge AI

Shopify is transitioning its mobile app development from React Native back to native Swift and Kotlin codebases, citing that AI agents can now handle much of the implementation, translation, testing, and review work.

Why it matters: Suggests that AI agents are becoming capable enough to influence fundamental mobile development strategies, impacting phone AI development.

Runtimes and quantization

b11012NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp adds Vulkan support for Qwen4exp hc ops and lists supported platforms including Apple Silicon, Intel, Linux, Windows, and Android with various backends like CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, and OpenCL.

Why it matters: Expands Qwen model support and hardware acceleration options for local inference across diverse devices and operating systems.

Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity BoundsNEW

arXiv · · Runtimes and quantization

DISCERN is a sequential two-tier protocol for certified paired risk-difference auditing of model updates, using a zero-label tier for benign updates and an audited tier for sampled disagreements.

Why it matters: Provides a method to ensure model updates do not degrade performance, which is relevant for maintaining local model quality.

Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUsNEW

arXiv · · Runtimes and quantization

The Tri-Metric Router is a deterministic, training-free policy for RAG on commodity GPUs, addressing the Compression Paradox by adaptively selecting among Raw, Neural, and Lexical pipelines based on CPU-side signals.

Why it matters: Improves RAG efficiency and memory management for long contexts on commodity GPUs, beneficial for local AI setups.

Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV CachesNEW

arXiv · · Runtimes and quantization

Fathom is a key scan method for sparse decoding over offloaded KV caches, allowing each query to decide how many bits of each key channel to read, speeding up decoding for large models like Qwen3-8B.

Why it matters: Enhances decoding speed and efficiency for large models with offloaded KV caches, improving local inference performance.

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM InferenceNEW

arXiv · · Runtimes and quantization

ASPIRE is a non-synchronized batched self-speculative decoding framework for long-context LLM inference, using a unified mixed forward and a lightweight online speculation scheduler.

Why it matters: Improves efficiency and throughput for long-context LLM inference, beneficial for local and sandbox environments.

proto-v0.3.0NEW

vLLM releases · · Runtimes and quantization

vllm-proto 0.3.0 has been released.

Why it matters: Indicates ongoing development in vLLM related tools, which are used for efficient LLM serving.

v0.34.2NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.2 has been released, incorporating updates from llama.cpp.

Why it matters: Indicates continuous improvement and integration of core runtime optimizations for local model serving.

b11009NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp includes a fix for split state and granularity for fused QKV in gemma4 and qwen35 models, and handles fused full attention layers for qwen35/qwen35moe.

Why it matters: Improves support and efficiency for specific Gemma and Qwen models within llama.cpp, enhancing local inference capabilities.

Agent sandboxes (E2B and peers)

Locating Hidden Failures Makes Long-Horizon Agents More ReliableNEW

arXiv · · Agent sandboxes (E2B and peers)

A study of 2518 agent trajectories across software engineering, computer use, and science classified 6967 mistakes into 78 failure types, noting agents often fail to recover.

Why it matters: Highlights reliability challenges in long-horizon agents, informing the design and testing of agent sandboxes.

A Large-Scale Empirical Study of Quality Assurance Practices and Gaps in AI AgentsNEW

arXiv · · Agent sandboxes (E2B and peers)

A large-scale empirical study of QA practices in 157 open-source LLM-based agent projects found that current QA primarily focuses on basic functionality and high-risk actions, with fragmented coverage.

Why it matters: Identifies gaps in quality assurance for AI agents, informing better development and testing practices for agent sandboxes.

How Lark Uses E2B to Safely Test Apps with Customer DataNEW

E2B Blog · · Agent sandboxes (E2B and peers)

Lark uses E2B sandboxes to safely test customer applications within Docker-based development environments.

Why it matters: Demonstrates a practical application of E2B sandboxes for secure testing, relevant for agent development.

e2b@2.50.0NEW

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK 2.50.0 removes V1 template build operations and schemas from generated API clients, with template builds now going through the Template SDK.

Why it matters: Indicates an evolution in E2B's API and SDK for managing sandboxes, requiring developers to adapt to new template management.

e2b@2.49.1

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK 2.49.1 adds retry logic for control-plane HTTP requests after 429 responses and clarifies behavior for onResume/on_resume options.

Why it matters: Improves the robustness and reliability of E2B sandbox interactions, beneficial for agent development.

2026.30

E2B infra releases · · Agent sandboxes (E2B and peers)

E2B infrastructure release 2026.30 removes deprecated access-token authentication, adds sandbox-list sorting and filtering, workload identity configuration, and feature-gated secrets operations.

Why it matters: Enhances security, management, and functionality of E2B sandboxes, providing more robust environments for agent development.

Open models for local use

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge PanelsNEW

arXiv · · Open models for local use

A study on LLM-as-judge panels found a positive same-family lift (3.4-8.4 percentage points) in preference, with a corrected estimator comparing judges while holding candidate family fixed.

Why it matters: Reveals biases in LLM-as-judge evaluations, which is important for understanding model performance and selection in local AI development.

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive