K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

Local AI Runtimes Advance, Sandboxes Enhance Agent Control, Mobile AI Shifts

Published , 04:44 Bogota (UTC-5) · 29 sourced items, 21 new since the previous edition · Read the foundations review · RSS

Today in 5 points

Laptop AI (MacBook, MLX)

GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three SystemsNEW

arXiv · · Laptop AI (MacBook, MLX)

This paper predicts single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors. It scores 318 phase-depth measurements from 53 host-file configurations on Apple M4 Max systems and an NVIDIA RTX

Why it matters: It offers a method to predict llama.cpp throughput on different local hardware configurations (Apple Silicon, NVIDIA), which helps users choose optimal models and quantization for

Desk-side boxes (Mac Studio, DGX Spark, OEM)

TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOpsNEW

arXiv · · Desk-side boxes (Mac Studio, DGX Spark, OEM)

TriCalRAG is a benchmark for evaluating open-weight LLMs served locally via vLLM on a single high-memory workstation GPU (NVIDIA RTX PRO 6000, 96GB) for root cause analysis in AIOps. It evaluates Qwen2.5-14B and Mistral-Small under zero-shot, few-shot, and RAG

Why it matters: This benchmark provides a way to assess the performance of local LLMs for critical enterprise tasks like AIOps, considering data privacy, latency, and cost benefits of on-premise d

Phone and edge AI

Beyond Noise: Understanding and Overcoming Temperature Effects in Analog DNN InferenceNEW

arXiv · · Phone and edge AI

This study investigates the impact of temperature on analog DNN inference, characterizing stochastic and systematic non-idealities across operating temperatures. It compares simulation-based and hardware-based mitigation strategies to improve robustness.

Why it matters: Understanding and mitigating temperature effects is crucial for deploying energy-efficient analog accelerators on constrained platforms like mobile and embedded devices, ensuring r

Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-ExecutionNEW

arXiv · · Phone and edge AI

This paper considers inter-operator and intra-operator parallelism for mobile heterogeneous inference, defining partition-aware DAG scheduling. The best strategy depends on the inference DAG structure, motivating a joint formulation for operator partition choi

Why it matters: It aims to improve inference latency on mobile heterogeneous platforms by optimizing how tasks are distributed and executed across CPUs and GPUs.

Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real RobotsNEW

arXiv · · Phone and edge AI

This paper investigates the real-world applicability of decentralized multi-agent reinforcement learning (MARL) for multi-robot multi-machine tending. It proposes Feature-fusion Multi-Agent Proximal Policy Optimization (FMAPPO) for safe decentralized task assi

Why it matters: It addresses challenges in deploying multi-agent RL on physical multi-robot systems in industrial environments, which is relevant for edge AI applications.

A Game-Theoretic Framework for Incentive-Compatible AI training Under Renewable-Energy ConstraintsNEW

arXiv · · Phone and edge AI

A game-theoretic model is developed for carbon-aware AI training, where autonomous agents strategically choose participation and training intensity under limited renewable energy. Agents balance learning returns, rewards for green-energy budgets, and penalties

Why it matters: This framework addresses the energy footprint of distributed and collaborative AI training, especially on edge devices, by aligning computational workloads with renewable energy su

Native is now the future of mobile at Shopify

Simon Willison · · Phone and edge AI

Shopify is moving from React Native back to separate Swift and Kotlin codebases for their native apps. This decision is attributed to agents now being able to handle enough implementation, translation, testing, and review work to reduce the cost of maintaining

Why it matters: This indicates that AI agents are becoming capable enough to reduce the development overhead of native mobile app development, potentially influencing how AI-powered features are i

Quoting Calif Research

Simon Willison · · Phone and edge AI

Calif Research developed a zero-click worm, WeWorm, that spreads through WeChat calls across iOS and Android. An AI team found the bug and wrote the remote code execution exploit in about two days, with the worm taking one more week to build.

Why it matters: This demonstrates the accelerated capability of AI in security research, enabling rapid discovery and exploitation of vulnerabilities on mobile platforms, highlighting potential ri

v0.17.0

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.0 optimizes local attention for longer contexts and adds Metal residency support for Apple Silicon acceleration. It extends Gemma 4 (12B) with multimodal capabilities, multi-token prediction acceleration, and extended context window support.

Why it matters: This update improves performance and capabilities for running LLMs like Gemma 4 on Apple devices and other mobile platforms, enabling longer contexts and multimodal features.

Runtimes and quantization

v0.34.1NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.1 includes fixes for the ChatGPT model selector spacing, MLX runner prefix cache eviction, and system free memory checks before loading MLX models. It also raises the token repeat limit, scopes MLX array lifetimes, keeps gemma3n projector off the

Why it matters: This update improves memory management, token handling, and general stability for local AI model execution, especially on MLX-compatible hardware.

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive PredictionNEW

arXiv · · Runtimes and quantization

This paper explores the relationship between learned memory and data in autoregressive prediction, showing they are governed by a predictive-energy spectrum. It introduces a minimax law for prediction blocks and learned state values, where data sets resolution

Why it matters: It provides theoretical insights into how data and memory resources interact in autoregressive models, which can inform efficient model design and deployment.

An Efficient and Modular Framework for Targeted Harm Mitigation in LLMSNEW

arXiv · · Runtimes and quantization

A modular correction framework is proposed to mitigate harms in LLMs by augmenting pretrained models with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism. This allows expert adapters to activate mid-sequence without invalidating the KV ca

Why it matters: This offers a flexible and scalable method for improving LLM alignment and safety locally, without costly retraining or tight coupling to the base model.

Hardware-Aware Learned Representation Compression for Distributed In-Sensor VisionNEW

arXiv · · Runtimes and quantization

OASIS is a distributed in-sensor vision framework that uses a lightweight encoder to generate compact, task-relevant representations before off-chip transmission. It supports 4-bit quantization and Huffman coding for spatial structures, and Sobol-based hyperdi

Why it matters: This framework addresses compute and memory constraints on integrated logic chips in CMOS image sensors, enabling efficient early-stage processing for on-device vision tasks.

Entropy-Punctured Bloom Filters for Memory-Efficient Machine LearningNEW

arXiv · · Runtimes and quantization

Entropy-punctured Bloom Filters are proposed as a memory-aware encoding strategy for machine learning. This approach removes low-variability bit positions from fixed-length Bloom Filter encodings to produce reduced representations that preserve predictive stru

Why it matters: It offers a method for creating compact feature representations, which is important for machine learning on constrained platforms due to storage, transmission, bandwidth, or privac

b10975: cuda : enable i16 and i32 for DUP (#28897)NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp updates enable i16 and i32 for DUP on CUDA.

Why it matters: This expands data type support for DUP operations on CUDA, potentially improving performance or compatibility for certain models in llama.cpp.

v0.34.1-rc1NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.1-rc1 adds an MLX patch to the Docker build context.

Why it matters: This update ensures MLX-specific optimizations are included in Docker builds for Ollama, benefiting users deploying on Apple Silicon.

v0.4.1NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp 0.4.1 adds support for Maple 20B-A1B, Tencent Hy 4, and Spark2.5 models. It improves JSON schema handling, chat parsing, logging, and server child-process management. Core changes include Kimi-K3 recurrent-state rollback support and fixes for MTP con

Why it matters: This release expands the range of models supported by llama.cpp and improves its robustness and usability for local inference, especially with new architectures and server deployme

Agent sandboxes (E2B and peers)

CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real TimeNEW

arXiv · · Agent sandboxes (E2B and peers)

CoArena is a real-time evaluation system for computer-use and multi-agent systems. It measures agent use directly by having real users submit tasks, two systems execute concurrently in sandboxed desktops, and users judge outcomes for a public leaderboard.

Why it matters: This provides a dynamic and user-driven method for evaluating agents in sandboxes, addressing the limitations of static benchmarks that can become outdated or leak into training da

CodeTS: Verifiable Text-to-Time Series Generation via Executable CodeNEW

arXiv · · Agent sandboxes (E2B and peers)

CodeTS is a verifiable framework for Text-to-Time Series Generation that uses code as an intermediate interface. It reformulates the process as Text-to-Code-to-TS, mapping textual descriptions into an explicit code space for time series synthesis through code

Why it matters: This framework offers a verifiable and explicit method for generating time series from natural language, which can be useful for agents operating in sandboxes that require precise

OpenAl4S: Code as Action, Science as SessionsNEW

arXiv · · Agent sandboxes (E2B and peers)

OpenAI4S is an open-source scientific research agent built around "Code as Action, Science as Sessions." It combines a persistent computing runtime with research-session management, using structured tool calls and code cells executed in persistent Python and R

Why it matters: It provides a framework for inspectable, resumable, and reproducible AI co-scientist workflows in sandboxes, preserving computational state and provenance.

e2b@2.49.1

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK v2.49.1 rejects invalid Sandbox.create lifecycle options and Sandbox.connect onResume values before requiring an API key. It documents that onResume needs a control plane that knows the option and retries control-plane HTTP requests up to three times a

Why it matters: This update improves the robustness and error handling of the E2B SDK for agent sandboxes, ensuring more reliable connections and lifecycle management.

2026.30

E2B infra releases · · Agent sandboxes (E2B and peers)

E2B infra release 2026.30 removes deprecated access-token authentication, adds sandbox-list sorting and filtering, and includes sandbox workload identity configuration. It also adds feature-gated /secrets operations and dynamic log routing.

Why it matters: This infrastructure update enhances security, management, and operational capabilities for E2B sandboxes, providing more control and flexibility for agent deployments.

Build an Agent Workbench on OpenAI's Agents API

E2B Blog · · Agent sandboxes (E2B and peers)

A workbench can be built on OpenAI's Agents API (beta) and E2B sandboxes, featuring application-managed lifecycle, one sandbox per chat, and pause and fork capabilities.

Why it matters: This provides a blueprint for developing sophisticated agent sandboxes with advanced lifecycle management, useful for local AI development and testing.

e2b@2.49.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK v2.49.0 exposes a configurable minimum free-disk target with minFreeDiskMb in JavaScript and min_free_disk_mb in Python.

Why it matters: This allows users to manage disk space more effectively within E2B sandboxes, which is important for agents that generate or process large amounts of data.

Open models for local use

Calibrating Interpretability Instruments Before Trusting Their VerdictsNEW

arXiv · · Open models for local use

This note documents six failures of causal interpretability measurements for LLM internals, where instruments can return plausible numbers instead of errors. Examples include covariance-matched nulls saturating, per-head attribution overshooting, and interchan

Why it matters: It highlights the importance of calibrating interpretability instruments to avoid misinterpreting LLM internal workings, which is crucial for understanding and debugging local mode

Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model FamiliesNEW

arXiv · · Open models for local use

Alignment applied after pretraining is shown to be shallow, as a single direction in a model's residual stream can be edited to remove refusal of harmful requests. Moral comprehension is native to pretraining, while the refusal gate is a post-training construc

Why it matters: This research provides insight into how refusal mechanisms in LLMs operate, suggesting that core moral comprehension is distinct from the refusal gate, which can be relevant for fi

The Misery of Mechanistic Interpretability: A Formal PerspectiveNEW

arXiv · · Open models for local use

This paper proposes a formal verification framework for the faithfulness of interpretable replacement networks (IRNs) used to understand language models. It shows that minor input perturbations can flip dominant IRN features across five open-weight model famil

Why it matters: It addresses the challenge of trusting mechanistic interpretability methods, which is important for developers trying to understand and debug the behavior of local LLMs.

Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated RunsNEW

arXiv · · Open models for local use

This study introduces "same-input rerun" to measure action-level divergence in clinical LLM agents, where identical inputs can lead to materially different orders despite benchmarks reporting the same verdict. It applies six reliability metrics to MedAgentBenc

Why it matters: It highlights a critical reliability issue for agents, especially in sensitive applications, where consistent behavior from local LLM agents under identical conditions is essential

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive