K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

Local AI Runtimes Update, Phone AI Advances, Agent Sandboxes Streamlined

Published , 04:44 Bogota (UTC-5) · 26 sourced items, 18 new since the previous edition · Read the foundations review · RSS

Today in 5 points

K/20X research paper

Laptop AI (MacBook, MLX)

LOIP:Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge DevicesNEW

arXiv · · Laptop AI (MacBook, MLX)

LOIP is a collaborative lossless LLM inference system for edge devices that uses offloading-based interleaved pipeline parallelism. It aims to overlap model offloading with computing and communicating to address memory budgets and fluctuating network bandwidth

Why it matters: This system provides lossless LLM inference on edge devices, including laptops, by optimizing offloading and parallelism, which is important for local AI performance.

Phone and edge AI

Extending FunctionGemma for Practical On-Device Mobile Function CallingNEW

arXiv · · Phone and edge AI

Research extends FunctionGemma 270M-it for Android function calling using MOBILEACTIONSEXTENDED, a synthetic dataset of ~9,500 conversations across fifteen device-control categories. Fine-tuning improves end-to-end accuracy on this dataset from 29.3% for the b

Why it matters: This work improves on-device function calling for Android, enabling local AI assistants to map natural language to system actions.

On the Effect of Bit-Level Parameter Perturbations in Machine Learning and Deep Learning ModelsNEW

arXiv · · Phone and edge AI

Research investigates the effect of bit-level parameter perturbations in classical machine learning models (Hidden Markov Models, Support Vector Machines) and deep learning models (Multilayer Perceptrons, Long Short-Term Memory networks). Applied to the Drebin

Why it matters: This study highlights the sensitivity of machine learning models, including those potentially used on phones, to small parameter changes, which is relevant for model robustness and

Learning from Humans for Proactive Assistance in Human-Robot Collaborative TransportNEW

arXiv · · Phone and edge AI

PROACT is a framework for human-robot collaborative transport that integrates predictions of human collaborative behavior with compliant robot control. This framework addresses the challenge of a robot acting as an effective partner in tasks like relocating la

Why it matters: This framework is relevant for edge AI applications where robots need to interact proactively and compliantly with humans, potentially using on-device intelligence.

HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive LearningNEW

arXiv · · Phone and edge AI

HYDRA (Hybrid Drift Adaptation) is a proactive adaptation framework for Android malware detection that learns drift-invariant representations from hierarchically structured data. It models applications using a hybrid graph structure combining Control Flow Grap

Why it matters: This framework offers a proactive approach to maintaining the performance of machine learning detectors for Android malware, which is relevant for phone AI security.

datasette-auth-github 1.0

Simon Willison · · Phone and edge AI

datasette-auth-github 1.0, a GitHub login plugin, was released. It fixes an issue where authenticated sessions were not lasting long due to cookies lacking a Max-Age parameter.

Why it matters: This plugin update improves session persistence for web applications, which can be relevant for agents or services running on mobile devices or accessed from them.

v0.17.1

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.1 introduces a bug fix to a tool call integer type issue.

Why it matters: This update improves the stability of tool calling in LiteRT-LM, which is important for reliable agentic behavior on phone AI.

Claude Cowork and chat are now one Claude

Simon Willison · · Phone and edge AI

Claude Cowork and chat are merging into one Claude, rolling out to Pro and Max plans on web, desktop, and mobile.

Why it matters: This consolidation simplifies the Claude offering, making it a general agent across platforms, including mobile, which is relevant for phone AI.

Runtimes and quantization

b11119NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp release b11119 reduces the size of the sampler probe. It supports macOS Apple Silicon (arm64), macOS Intel (x64), iOS, various Linux configurations (x64, arm64, s390x, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL FP32/FP16, Snapdragon), Android arm6

Why it matters: This update to a local inference runtime offers broad platform support, including Apple Silicon, Intel, Linux, Android, and Windows, with various hardware accelerators.

Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer RefinementNEW

arXiv · · Runtimes and quantization

Cross-layer Activation-aware Sensitivity Allocation (CASA) is proposed for mixed-precision LLM quantization. It addresses limitations of scalar sensitivity proxies, which can incur significant multiplicative distortion, by replacing the scalar proxy in Stage 1

Why it matters: This method aims to improve the accuracy and reliability of mixed-precision quantization for LLMs, which can reduce memory and computational requirements for local inference.

FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware QuantizationNEW

arXiv · · Runtimes and quantization

FuncCode is a basis-agnostic compression approach for Kolmogorov-Arnold Networks (KANs) that forms shared codebooks from sampled edge responses. It codes basis and base branches independently and exports quantized, bit-packed codebooks and per-edge indices. Sa

Why it matters: This compression method for KANs can reduce parameter memory, making these models more feasible for local deployment on devices with limited resources.

Component Type, Not Reconstruction Error, Predicts Attention Quantization SensitivityNEW

arXiv · · Runtimes and quantization

A study on attention quantization sensitivity across nine open-weight language models (1.3B-8B parameters) found that within a component type (Q, K, V, or O), reconstruction error explains less than 10% of the variance in perplexity sensitivity. It quantizes o

Why it matters: This research indicates that component type, rather than reconstruction error, is a better predictor of attention quantization sensitivity, which is relevant for optimizing local L

CompKV: Compensation-Aware KV Selection for Long-Context LLM InferenceNEW

arXiv · · Runtimes and quantization

CompKV is introduced as a compensation-aware sparse attention framework for long-context LLM inference. It addresses the limitation of existing sparse attention methods that select tokens based on attention mass and then compensate, by explicitly optimizing se

Why it matters: This framework aims to improve long-context LLM inference by optimizing KV cache memory traffic, which is crucial for running larger models locally with extended contexts.

v0.34.4NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.4 fixes intermittent "model not found" errors, applies structured outputs in a single pass on thinking models, and avoids System Events for ChatGPT/Codex detection. It includes llama.cpp and MLX version updates, dynamic Gemma 4 image resolution se

Why it matters: This Ollama update improves stability, performance, and feature support for local LLM inference, including specific optimizations for Gemma and Qwen models on MLX.

b11115NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp release b11115 adds an OpenCL binary kernel (kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin) and an A8 Q4_K non-MoE dp4a binary kernel. It also renames binary kernel selection helpers.

Why it matters: This update enhances OpenCL support in llama.cpp, potentially improving performance for users with compatible GPUs for local inference.

v0.34.3NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.3 now advertises each model's thinking controls and default via GET /api/show. Nemotron H vision models are supported on Apple Silicon with MLX. The macOS app no longer reopens closed windows upon activation.

Why it matters: This Ollama update improves model introspection, adds support for Nemotron H vision models on Apple Silicon, and refines the macOS user experience for local AI users.

Agent sandboxes (E2B and peers)

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel PlanningNEW

arXiv · · Agent sandboxes (E2B and peers)

GroupTravelBench is introduced as a benchmark for multi-user, multi-turn travel planning for LLM agents. It comprises 650 tasks across three difficulty levels, built from real user profiles, POI data, and ticket prices, running in a synchronous group-chat sand

Why it matters: This benchmark addresses the complexity of multi-user planning for LLM agents, moving beyond single-user scenarios and providing a reproducible environment for evaluation in sandbo

e2b@2.51.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK e2b@2.51.0 removes SDK-side defaults from API request payloads for sandbox create/fork/connect, allowing API defaults to apply. Sandbox create and connect now use v2 API endpoints, which default timeout to 5 minutes and secure envd access. The secure o

Why it matters: This SDK update streamlines sandbox creation and connection, ensuring API defaults are used and all sandboxes are secured, which is important for agent development.

How Lark Uses E2B to Safely Test Apps with Customer Data

E2B Blog · · Agent sandboxes (E2B and peers)

Lark uses E2B sandboxes to test customer applications in Docker-based development environments.

Why it matters: This demonstrates a practical application of E2B sandboxes for safely testing applications with customer data, relevant for agent development and deployment.

e2b@2.50.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK e2b@2.50.0 removed V1 template build operations and schemas from generated API clients, as the API no longer serves them. Template builds now go through the Template SDK.

Why it matters: This SDK update streamlines template build processes by consolidating them under the Template SDK, affecting how agent sandboxes are configured.

Open models for local use

Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and TrainingNEW

arXiv · · Open models for local use

Hill Sampling is introduced as a method to improve LLM solutions for verifiable problems by repeatedly sampling candidate program edits from a frozen LLM, retaining the best, and conditioning subsequent samples on it. It sets a new state of the art on circle p

Why it matters: This method offers a simple way to improve LLM performance on algorithmic problems at test time without elaborate search harnesses or model parameter updates.

You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMsNEW

arXiv · · Open models for local use

An empirical study of dynamic expert pruning in fine-grained Mixture-of-Experts (MoE) LLMs finds that expert selection redundancy and pruning method effectiveness are central questions. The study covers twelve MoE checkpoints across nine architecture families

Why it matters: This research investigates how to make fine-grained MoE LLMs cheaper for inference by understanding and exploiting redundancy in expert selection.

Same Quantity, Different Answer: Numerical Representation Invariance in Language ModelsNEW

arXiv · · Open models for local use

A study on numerical representation invariance in language models evaluates five open-weight systems on 3,600 exact-rational problems and 8,600 prompts with identity-preserving transformations. Canonical accuracy is high (0.969-0.996), but orbit correctness an

Why it matters: This research highlights challenges in LLM numerical reasoning across different representations, which is important for reliable local AI applications requiring precise calculation

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-MakingNEW

arXiv · · Open models for local use

DFAH-Bench operationalizes the Determinism-Faithfulness Assurance Harness (DFAH) to benchmark observable agent instability in financial decision-making. It pairs decision agreement with tool-path agreement on qualified replays. Decision agreement is 94.2-95.1%

Why it matters: This benchmark helps evaluate the stability and faithfulness of AI agents, particularly in critical applications like financial decision-making, which is relevant for agent sandbox

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive