K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

Local AI Runtimes, Phone AI Efficiency, and Agent Sandboxes See Updates

Published , 04:44 Bogota (UTC-5) · 24 sourced items, 15 new since the previous edition · Read the foundations review · RSS

Today in 5 points

K/20X research paper

Phone and edge AI

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV CachingNEW

arXiv · · Phone and edge AI

TierKV is a mobile LLM inference framework that uses Predictive Multi-Tier Cache Optimization (PMCO) to manage KV cache memory, predicting demand and assigning tokens to different tiers.

Why it matters: Mitigates the KV cache memory bottleneck for long-context LLMs on mobile devices, enabling more capable on-device AI.

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM InferenceNEW

arXiv · · Phone and edge AI

SpecQuant is a training-free framework that combines speculative decoding with multi-parent quantization for adaptive LLM inference, using dynamically routed quantized variants.

Why it matters: Offers substantial speed boosts for LLMs on consumer hardware by adapting inference based on task complexity and available precision.

The Weight Is Over - Interactive Diffusion on Consumer GPUsNEW

arXiv · · Phone and edge AI

Research on interactive diffusion on consumer GPUs introduces an embedding translator to reduce weight and latency, a sweep recipe for diffusion pipelines, and an on-device image generation editor.

Why it matters: Addresses memory and latency challenges for diffusion models on consumer hardware, enabling interactive on-device image generation.

Do small language models know what they don't know?NEW

arXiv · · Phone and edge AI

Research on small language models (SLMs) finds token-level entropy ineffective for confidence signals, but semantic entropy, derived from clustering multiple samples, provides a viable signal.

Why it matters: Helps improve the reliability and accuracy of SLMs on consumer hardware by enabling them to signal uncertainty.

datasette-auth-github 1.0

Simon Willison · · Phone and edge AI

datasette-auth-github 1.0 was released, fixing an issue where authenticated sessions expired quickly due to missing cookie parameters.

Why it matters: Improves user experience and session persistence for applications using this plugin, potentially relevant for agent sandboxes or local web UIs.

v0.17.1

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.1 includes a bug fix for a tool call integer type issue.

Why it matters: Ensures correct functionality for LiteRT-LM, which is relevant for phone AI and efficient LLM inference.

Claude Cowork and chat are now one Claude

Simon Willison · · Phone and edge AI

Claude Cowork and chat are merging into a single Claude experience, rolling out to Pro and Max plans on web, desktop, and mobile, indicating a move towards a general agent.

Why it matters: Reflects a trend towards unified AI agent platforms across devices, impacting how users interact with AI on phones and desktops.

Runtimes and quantization

b11070NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp's hexagon backend received an overhaul of buffer and DMA handling to support 64bit mappings, including extended buffer mappings and 64bit DMA in most binary operations.

Why it matters: Improves llama.cpp's performance and memory handling on hexagon targets, potentially enabling larger models or more efficient inference.

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context DecodingNEW

arXiv · · Runtimes and quantization

Elastic Threshold Attention (ETA) is a trainable architecture for long-context decoding that uses dynamic, contextual thresholds to manage context allocation, aiming for hardware-accelerated speed without quality loss.

Why it matters: Addresses memory-bandwidth bottlenecks in long-context decoding, potentially improving efficiency and performance for local LLM inference.

Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal ResidualsNEW

arXiv · · Runtimes and quantization

Research clarifies LLM quantization error components, distinguishing those addressed by weight compensation from those needing transformation design, offering guidelines for aggressive W4A4 quantization.

Why it matters: Improves understanding and implementation of aggressive quantization (W4A4), crucial for reducing memory and inference costs of LLMs locally.

Multi-Domain Clustering via Measure QuantizationNEW

arXiv · · Runtimes and quantization

A general framework for multi-domain clustering via measure quantization learns shared cluster prototypes across multiple domains, using a mini-batch optimization strategy for scalability.

Why it matters: Provides a scalable method for data analysis, which could be relevant for optimizing data processing for local AI models or training.

RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative DecodingNEW

arXiv · · Runtimes and quantization

RheoSampling addresses the "one-hot dilemma" in stochastic dynamic-tree speculative decoding, which causes a drop in acceptance rate by collapsing draft distributions into one-hot probabilities.

Why it matters: Improves the efficiency and effectiveness of speculative decoding for LLMs, particularly in stochastic sampling scenarios.

b11065NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp received CUDA tuning for FA on Gemma 4 models on Ampere or newer GPUs, with various build targets listed for different platforms and backends.

Why it matters: Improves performance for Gemma 4 models on NVIDIA GPUs and highlights llama.cpp's broad platform support for local inference.

v0.34.3

Ollama releases · · Runtimes and quantization

Ollama v0.34.3 adds support for Nemotron H vision models on Apple Silicon with MLX, advertises thinking controls via /api/show, and fixes model pulls from HuggingFace.

Why it matters: Expands model compatibility and improves functionality for Ollama users, especially those on Apple Silicon, for local AI inference.

Agent sandboxes (E2B and peers)

e2b@2.51.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK @2.51.0 removes SDK-side defaults from API request payloads, ensuring API defaults apply. Sandbox create/connect now use v2 API endpoints, securing envd access and defaulting timeout to 5 minutes.

Why it matters: Streamlines sandbox configuration and enhances security for agent sandboxes by ensuring API defaults are used and connections are secured.

How Lark Uses E2B to Safely Test Apps with Customer Data

E2B Blog · · Agent sandboxes (E2B and peers)

Lark uses E2B sandboxes for testing customer applications within Docker-based development environments.

Why it matters: Demonstrates a practical application of E2B sandboxes for secure and isolated testing, relevant for agent development.

e2b@2.50.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK @2.50.0 removed V1 template build operations and schemas from API clients, with template builds now handled via the Template SDK.

Why it matters: Streamlines the template build process for E2B sandboxes, simplifying development for agent sandboxes.

Open models for local use

ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read FactorNEW

arXiv · · Open models for local use

ServeGuard is a method to ensure verifiable, bounded-residual confinement of operator-invisible channels in third-party adapters for open-weight language models, aiming to prevent backdoors.

Why it matters: Enhances the security and trustworthiness of open-weight LLMs by providing a way to verify that adapters do not hide malicious channels.

K2-V2: A 360-Open, Reasoning-Enhanced LLMNEW

arXiv · · Open models for local use

K2-V2 is a 360-open LLM designed for reasoning adaptation, conversation, and knowledge retrieval, described as a strong fully open model.

Why it matters: Offers a powerful, open-source base model for local deployment and fine-tuning, potentially rivaling larger models in performance.

Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30NEW

arXiv · · Open models for local use

Research explores steering open-weight LLM responses towards moral foundations using prompt-level persona steering and activation-level ActAdd interventions.

Why it matters: Provides insights into controlling LLM behavior and alignment, which is relevant for deploying models responsibly in local or sandbox environments.

Do Personality-Tuned LLMs Make Better Social Agents?NEW

arXiv · · Open models for local use

Research investigates if personality-aware fine-tuning improves consistency and controllability of personality-conditioned dialogue generation in LLMs, fine-tuning Qwen2.5-7B-Instruct and Ministral-8B-Instruct.

Why it matters: Enhances the ability to create more consistent and controllable social agents using local LLMs, improving their utility in simulations.

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive