K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

Cross-Platform AI Inference Improves, Agent Sandboxes Gain Features, Security Concerns Rise

Published , 04:44 Bogota (UTC-5) · 26 sourced items, 17 new since the previous edition · Read the foundations review · RSS

Today in 5 points

Phone and edge AI

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-TuningNEW

arXiv · · Phone and edge AI

An efficient method distills reasoning capabilities into compact video-language models (VLMs) for VideoQA. A 2B-parameter model, fine-tuned with ~900 uncertainty-selected examples and synthetic Chain-of-Thought (CoT) rationales, outperformed VLMs up to 4x larg

Why it matters: This demonstrates efficient training for smaller VLMs, making advanced reasoning capabilities more accessible for phone AI and other resource-constrained local deployments.

Latent Undertow: How Ordinary Typos Break ProbesNEW

arXiv · · Phone and edge AI

Ordinary typos in LLM inputs, while not affecting user intent or model response, significantly perturb hidden states. Probes detecting malicious prompts by reading hidden states show a 43-56 degree rotation at the perturbed token.

Why it matters: This highlights a vulnerability in probe-based safety mechanisms to common typos, impacting the reliability of AI safety tools, especially for phone AI and agent sandboxes.

Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive AcceleratorsNEW

arXiv · · Phone and edge AI

A cross-platform study on INT8 quantized inference across seven hardware classes found that INT8 portability fails on three axes. The sign of INT8 speedup depends on the CPU's dot-product ISA.

Why it matters: This study challenges the assumption of INT8 portability, providing critical insights for deploying quantized models on diverse edge and phone AI hardware, emphasizing ISA-specific

Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability PreservationNEW

arXiv · · Phone and edge AI

CRN v2 is a lightweight logit-level correction module (~34M trainable parameters) designed to fix errors in a frozen Gemma 4 E2B model's outputs without degrading base capabilities.

Why it matters: This offers a method to improve the accuracy of frozen models like Gemma 4 E2B without retraining the base, which is valuable for efficient deployment on phone AI and local devices

Native is now the future of mobile at Shopify

Simon Willison · · Phone and edge AI

Shopify is moving from React Native back to separate Swift and Kotlin codebases for their native mobile apps. AI agents can now handle enough implementation, translation, testing, and review work.

Why it matters: This demonstrates how AI agents are changing development paradigms, making native app development more feasible, which could impact how AI features are integrated into phone AI app

Quoting Calif Research

Simon Willison · · Phone and edge AI

Calif Research released a demo of WeWorm, a zero-click worm spreading through WeChat calls on iOS and Android. AI helped their team find the bug and write the first remote code execution exploit in about two days.

Why it matters: This highlights the accelerated pace of exploit development with AI, posing new security challenges for phone AI and local devices.

v0.17.0

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.0 features optimized Local Attention for reduced memory and longer contexts, Apple Silicon acceleration with Metal residency, extended Gemma 4 (12B) with multimodal capabilities, and multi-token prediction acceleration.

Why it matters: This update significantly improves performance and capabilities for Gemma 4 on Apple Silicon, directly impacting MacBook/MLX and phone AI users.

Runtimes and quantization

b10992NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp's llama-bench now supports a `--version` flag to print build info. The release lists extensive platform support including macOS Apple Silicon, Intel, iOS, various Linux configurations, Android, and Windows.

Why it matters: This update details broad platform support for llama.cpp, crucial for local AI inference across diverse hardware, from Apple Silicon to various Linux and Windows setups.

LLM Inference in a Flash!NEW

arXiv · · Runtimes and quantization

LLM inference faces challenges with longer sequences and heavier inference demands, compounded by memory and bandwidth not scaling as fast. Compute-in-Flash is presented as a solution to address memory bandwidth by moving computation closer to SSD memory.

Why it matters: This research explores solutions for memory bandwidth limitations in LLM inference, relevant for optimizing local AI performance, especially with large models and long contexts.

Skeletal Prototypes on Iterative Nerve ExpansionsNEW

arXiv · · Runtimes and quantization

Skeletal Prototypes on Iterative Nerve Expansions (SPINE) is a new prototype reduction method. It models each class as an embedded 1-complex, using a class-conditional Mapper graph for initial edge sets and fitting vertices under a classification objective.

Why it matters: This introduces a new method for prototype reduction, which can lead to more efficient model representations and potentially faster inference for local AI applications.

Adaptive Bayesian Partner Selection for Federated Clinical CentersNEW

arXiv · · Runtimes and quantization

Adaptive Bayesian Partner Selection (ABPS) is a peer-to-peer framework for federated learning in healthcare. It addresses heterogeneity and temporal concept drift by allowing centers to maintain a Beta-Bernoulli posterior over peers' utility.

Why it matters: This framework optimizes federated learning collaborations, which could be relevant for distributed AI training scenarios, potentially involving local devices.

LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural NetworksNEW

arXiv · · Runtimes and quantization

Latent Chip Adaptation from Probes (LCAP) is a population-informed framework for adapting Photonic Neural Networks (PNNs) to hardware variations. It learns a shared correction from historical chips and extracts a low-dimensional correction space.

Why it matters: This method addresses the sim-to-real gap for PNNs, which are efficient analog inference hardware. It could improve the reliability and deployment of specialized AI hardware.

b10991NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp's hexagon support added back missing contiguous fast-path and hvx_copy_uu for each run. The release lists extensive platform support, similar to i00.

Why it matters: This update improves performance for hexagon processors within llama.cpp, relevant for specific hardware deployments, potentially including some phone AI NPUs.

v0.34.2NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.2 includes llama.cpp updates.

Why it matters: Ollama integrating llama.cpp updates means improvements in llama.cpp are directly available to Ollama users, impacting local AI inference.

v0.34.1-rc1

Ollama releases · · Runtimes and quantization

Ollama v0.34.1-rc1 added an MLX patch to the docker build context.

Why it matters: This indicates Ollama's ongoing support and optimization for Apple Silicon (MLX), which is important for MacBook/Mac Studio users running local AI.

Agent sandboxes (E2B and peers)

Spurious Tool Use: When RL Agents Learn the Wrong Reason to ActNEW

arXiv · · Agent sandboxes (E2B and peers)

RL-trained LLM agents can learn shortcut tool-selection policies, invoking tools based on superficial prompt cues rather than task requirements. Controlled synthetic environments showed agents exhibiting substantial shortcut behavior.

Why it matters: This research identifies a risk in RL-trained agents learning spurious tool use, which is important for designing robust and reliable agents in sandboxes.

Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical MappingNEW

arXiv · · Agent sandboxes (E2B and peers)

Autonomous Software Engineering Agents (SWE-Agents) struggle with Architecture 0 due to "Unknown Unknowns." Empirical analysis showed pure-text reasoning devolves into consensus or impossible fabrications.

Why it matters: This highlights challenges for SWE-Agents in complex design tasks and the risks of "Specification Gaming" in sandboxes, crucial for developing more capable and reliable agents.

e2b@2.49.1

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK v2.49.1 rejects invalid Sandbox.create lifecycle options and Sandbox.connect onResume values before requiring an API key. It also retries control-plane HTTP requests up to three times after 429 responses.

Why it matters: These are stability and error handling improvements for the E2B sandbox, making agent development and deployment more robust.

2026.30

E2B infra releases · · Agent sandboxes (E2B and peers)

E2B infra release 2026.30 removes deprecated access-token authentication, adds sandbox-list sorting and filtering, adds sandbox workload identity configuration, and feature-gated /secrets operations.

Why it matters: These are significant infrastructure updates for E2B, enhancing security, management, and functionality for agent sandboxes.

Build an Agent Workbench on OpenAI's Agents API

E2B Blog · · Agent sandboxes (E2B and peers)

A workbench can be built on OpenAI's Agents API (beta) and E2B sandboxes, offering application-managed lifecycle, one sandbox per chat, pause, and fork.

Why it matters: This shows how E2B sandboxes integrate with OpenAI's Agents API, providing a robust environment for developing and managing AI agents.

e2b@2.49.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK v2.49.0 exposes a configurable minimum free-disk target with `minFreeDiskMb` in JavaScript, `min_free_disk_mb` in Python, and `--min-free-disk-mb` in template create.

Why it matters: This provides more control over disk space management in E2B sandboxes, which is useful for managing resources for agent operations.

Open models for local use

Evaluating Open-Weight E-Commerce Agents with Environment-Grounded VerificationNEW

arXiv · · Open models for local use

A new e-commerce environment is proposed for evaluating open-weight agents. It uses a deterministic, reproducible setup with precommitted customer and trajectory parameters, recording assistant actions and environment state for detailed post-trial assessment.

Why it matters: This provides a structured method for evaluating open-weight e-commerce agents, important for developers building and testing AI agents in sandboxes.

Decoy Direction Optimization: A Post-Hoc Defense Against LLM AbliterationNEW

arXiv · · Open models for local use

Decoy Direction Optimization (DDO) is a post-hoc weight-editing defense against Refusal Feature Ablation (RFA) attacks on LLMs. DDO injects a high-magnitude, nonlinear decoy signal into MLP neurons to mislead contrastive estimators used by RFA.

Why it matters: This offers a fast, post-hoc defense against attacks that bypass LLM safety guardrails, important for maintaining the integrity and safety of locally deployed or sandbox-based mode

Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 ArchitecturesNEW

arXiv · · Open models for local use

Few-shot prompting can degrade language models, and the effect is task-dependent. A new metric, "content delta," is proposed to isolate representation changes caused by prompt content, by subtracting shifts caused by prompt length using random text controls.

Why it matters: This research helps understand few-shot prompting behavior in open-weight models, crucial for optimizing their performance and reliability in local AI applications.

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using AgentsNEW

arXiv · · Open models for local use

Blindspot is a new benchmark for trajectory-level safety calibration of long-horizon tool-using agents. It evaluates complete user-agent-environment trajectories through adaptive adversarial interaction, stateful tool execution, and execution-grounded adjudica

Why it matters: This benchmark is crucial for assessing the safety and refusal calibration of LLM agents, especially those operating in sandboxes with tool use and persistent state.

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

NVIDIA Technical Blog · · Open models for local use

Alibaba released Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, with configurable reasoning, servable on NVIDIA GB300 NVL72.

Why it matters: This announces a large open-weight model, Qwen3.8-Max, which is relevant for high-end AI hardware and potentially for local AI development if smaller versions become available.

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive