<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>K/20X LABS: Local AI and E2B daily research brief</title><link>https://k20x.com/ai_setup/daily/</link><atom:link href="https://k20x.com/ai_setup/feed.xml" rel="self" type="application/rss+xml"/><description>Sourced daily research on local AI hardware, phone AI and agent sandboxes.</description><language>en</language><lastBuildDate>Fri, 09 Oct 2026 06:44:01 +0000</lastBuildDate><item><title>How ClickUp Runs AI Code on Sensitive Data at Enterprise Scale</title><link>https://e2b.dev/customers/clickup</link><guid isPermaLink="false">https://e2b.dev/customers/clickup</guid><pubDate>Thu, 08 Oct 2026 00:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>ClickUp Brain² uses E2B sandboxes to run model-written code on sensitive workspace data, leveraging isolated microVMs, prebuilt templates, and the ability to launch thousands of sandboxes quickly.</description></item><item><title>b11514</title><link>https://github.com/ggml-org/llama.cpp/releases/tag/b11514</link><guid isPermaLink="false">https://github.com/ggml-org/llama.cpp/releases/tag/b11514</guid><pubDate>Thu, 08 Oct 2026 20:17:43 +0000</pubDate><category>Runtimes and quantization</category><description>llama.cpp release b11514 includes a Musa FWHT fix and lists broad platform support, including macOS, iOS, Linux (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Android (CPU, Snapdragon), and Windows (CPU, OpenCL, CUDA).</description></item><item><title>proto-v0.5.0</title><link>https://github.com/vllm-project/vllm/releases/tag/proto-v0.5.0</link><guid isPermaLink="false">https://github.com/vllm-project/vllm/releases/tag/proto-v0.5.0</guid><pubDate>Fri, 09 Oct 2026 03:11:40 +0000</pubDate><category>Runtimes and quantization</category><description>vLLM proto-v0.5.0 has been released.</description></item><item><title>v0.40.2</title><link>https://github.com/ollama/ollama/releases/tag/v0.40.2</link><guid isPermaLink="false">https://github.com/ollama/ollama/releases/tag/v0.40.2</guid><pubDate>Fri, 09 Oct 2026 03:37:30 +0000</pubDate><category>Runtimes and quantization</category><description>Ollama v0.40.2 upgrades models downloaded with earlier versions for better performance and llama.cpp compatibility, temporarily keeping original copies as backups for safe downgrading.</description></item><item><title>System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives</title><link>https://arxiv.org/abs/2607.09842</link><guid isPermaLink="false">https://arxiv.org/abs/2607.09842</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>Corrections are reported for previous findings on system-prompt conditioning and hidden-state geometry in four open-weight language models, clarifying the curvature statistic and the reported norm quantity.</description></item><item><title>Closed-loop evaluation of LLM agents for embedded software development</title><link>https://arxiv.org/abs/2610.11447</link><guid isPermaLink="false">https://arxiv.org/abs/2610.11447</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>A benchmark is presented for closed-loop evaluation of LLM agents for embedded software development, where agents must translate requirements, self-verify, and iterate until the required device behavior is achieved.</description></item><item><title>AI4Fire: Evaluating Large Language Models on Wildfire Tasks</title><link>https://arxiv.org/abs/2610.10946</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10946</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>AI4Fire evaluates large language models on wildfire tasks, comparing bare and grounded runs across various models and tasks, finding that grounding helps most when the addition carries the answer and simple rules are hard to beat.</description></item><item><title>Power Side-Channel Membership Inference Attack on Embedded Machine Learning</title><link>https://arxiv.org/abs/2610.10909</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10909</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>PSCMIA is a power side-channel membership inference attack against embedded machine learning models that infers membership directly from power traces without requiring prediction probabilities or labels.</description></item><item><title>Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception</title><link>https://arxiv.org/abs/2610.12445</link><guid isPermaLink="false">https://arxiv.org/abs/2610.12445</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>White-box deception detection via probes is scaled up for monitoring LLM agents, using a large deception dataset and a novel probe architecture to achieve high AUC in SHADE-Arena and detect introspective deception.</description></item><item><title>NanoProof: Open and Efficient Automated Theorem Proving in Lean 4</title><link>https://arxiv.org/abs/2610.11605</link><guid isPermaLink="false">https://arxiv.org/abs/2610.11605</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>NanoProof is introduced as an open and efficient automated theorem prover in Lean 4, with all training data, tooling, pipeline, and weights released, focusing on compute efficiency for accessible training and evaluation.</description></item><item><title>Dynamics as Code: On Model Compression via Dynamic System</title><link>https://arxiv.org/abs/2610.11115</link><guid isPermaLink="false">https://arxiv.org/abs/2610.11115</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>Research proves that a dynamic system paradigm can be used for model compression, where high-dimensional parameters are encoded by a trajectory index and recovered during decompression, distinct from other compression methods.</description></item><item><title>Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models</title><link>https://arxiv.org/abs/2610.10933</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10933</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>Spora is introduced as a new approach for spiking language models that jointly designs spike encodings and attention operators, using binary temporal weights and unipolar/bipolar binary spiking to improve performance with fewer time steps.</description></item><item><title>Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory</title><link>https://arxiv.org/abs/2610.10926</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10926</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>An adaptive multi-discriminator Wasserstein GAN (MD-WGAN) framework is introduced for resource-constrained Internet of Vehicles environments, integrating reinforcement learning and game theory to manage machine learning workloads.</description></item><item><title>KDFP: A first-principles approach to knowledge distillation in large language models</title><link>https://arxiv.org/abs/2610.10854</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10854</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>KDFP is presented as a novel methodology for white-box general knowledge distillation in large language models, developed through a first-principles approach to create efficient and private systems for edge devices.</description></item><item><title>Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction</title><link>https://arxiv.org/abs/2610.10641</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10641</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>Research explores using large language models for next-visit diagnosis prediction, highlighting that existing reinforcement learning rewards can concentrate on a few correct diagnoses, leaving others uncovered, and that LLM tokenizers can split ICD codes into </description></item><item><title>Self-Organization from Constrained Geometric Radiation</title><link>https://arxiv.org/abs/2610.10621</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10621</guid><pubDate>Fri, 09 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>Research introduces constraint-induced self-organization via geometric radiation in coupled metric evolution systems, revealing a four-stage cycle and a novel wedge-shaped attractor topology.</description></item><item><title>b11516: hexagon: fix IM2COL patch-embed DMA ring overflow (#30189)</title><link>https://github.com/ggml-org/llama.cpp/releases/tag/b11516</link><guid isPermaLink="false">https://github.com/ggml-org/llama.cpp/releases/tag/b11516</guid><pubDate>Fri, 09 Oct 2026 06:04:09 +0000</pubDate><category>Runtimes and quantization</category><description>A fix for a DMA ring overflow issue in the Hexagon IM2COL patch-embed kernel in llama.cpp prevents stale data from leaking into output when IC*KH exceeds 255. The solution involves retiring the oldest descriptor when the ring is full.</description></item><item><title>Introducing E2B Secrets</title><link>https://e2b.dev/resources/introducing-e2b-secrets</link><guid isPermaLink="false">https://e2b.dev/resources/introducing-e2b-secrets</guid><pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>E2B Secrets is introduced to keep credentials outside the sandbox and inject them into HTTPS request headers. This allows agents to call external services safely.</description></item><item><title>v0.40.1-rc0</title><link>https://github.com/ollama/ollama/releases/tag/v0.40.1-rc0</link><guid isPermaLink="false">https://github.com/ollama/ollama/releases/tag/v0.40.1-rc0</guid><pubDate>Wed, 07 Oct 2026 22:22:54 +0000</pubDate><category>Runtimes and quantization</category><description>Ollama v0.40.1-rc0 drops a carried Metal residency patch for MLX because it is now upstream.</description></item><item><title>v0.40.1</title><link>https://github.com/ollama/ollama/releases/tag/v0.40.1</link><guid isPermaLink="false">https://github.com/ollama/ollama/releases/tag/v0.40.1</guid><pubDate>Thu, 08 Oct 2026 02:09:43 +0000</pubDate><category>Runtimes and quantization</category><description>Ollama v0.40.1 includes fixes for clef head reads on Windows, manifest symlinks on Windows, and removes an account step from CLI onboarding. It also drops a Metal residency patch for MLX as it is now upstream.</description></item><item><title>LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure</title><link>https://arxiv.org/abs/2608.13545</link><guid isPermaLink="false">https://arxiv.org/abs/2608.13545</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>LittleLearner is a 5B-parameter LLM trained on LittleCurriculum, an 88B-token pretraining corpus tailored to U.S. elementary school material. This creates a developmentally restricted sandbox to study knowledge and skill acquisition in models.</description></item><item><title>The Implications of Linguistic Illegibility for LLM Security</title><link>https://arxiv.org/abs/2609.02852</link><guid isPermaLink="false">https://arxiv.org/abs/2609.02852</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>The concept of &quot;linguistic illegibility&quot; is introduced, referring to scenarios where an LLM&#x27;s linguistic outputs or features do not reliably represent its internal computation. This implies that security mechanisms relying on a model&#x27;s language artifacts may b</description></item><item><title>Reinforcement Learning for Code Optimization</title><link>https://arxiv.org/abs/2607.25970</link><guid isPermaLink="false">https://arxiv.org/abs/2607.25970</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>Research addresses challenges in applying reinforcement learning (RL) to code optimization, where execution time drives the reward. It proposes making execution time learnable through a calibrated sandbox, composing correctness and speed in the RL environment,</description></item><item><title>RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments</title><link>https://arxiv.org/abs/2610.10409</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10409</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>RobotWorld is introduced as a simulation testbed for benchmarking multimodal agents for robot use across 84 tasks. It evaluates agents&#x27; ability to translate instructions and observations into physical task execution through robot interfaces.</description></item><item><title>Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures</title><link>https://arxiv.org/abs/2610.10261</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10261</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>Research evaluates Small Language Models (SLMs) for reverse-engineering Machine Learning pipeline structures from source code. SLMs are assessed for their code understanding and classification abilities to extract ML pipeline stages.</description></item><item><title>Democratizing MoE inference on commodity GPUs with CoMoE</title><link>https://arxiv.org/abs/2610.09424</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09424</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Desk-side boxes (Mac Studio, DGX Spark, OEM)</category><description>CoMoE is a communication-efficient MoE inference system designed to democratize MoE inference on commodity GPUs. It addresses communication bottlenecks on consumer GPUs, which lack high-bandwidth P2P interconnects, through novel host-centric routing.</description></item><item><title>emg2face: Expressive Facial Animation with High-Density Surface EMG</title><link>https://arxiv.org/abs/2610.09304</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09304</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>emg2face demonstrates expressive facial animation using high-density surface electromyography (HD-sEMG) as a non-optical alternative to face capture. It measures 64 EMG channels from the forehead and side of the face to estimate 3D facial landmarks.</description></item><item><title>Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs</title><link>https://arxiv.org/abs/2610.09033</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09033</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>A Quad-State Safety Evaluation assesses open-weight LLMs on non-canonical inputs using the Adversarial Surface-Form Robustness Dataset (ASRD). It finds that emoji and invisible Unicode variations cause almost no comprehension failure, with specific harmful com</description></item><item><title>How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis</title><link>https://arxiv.org/abs/2610.09000</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09000</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>Research investigates the fragility of on-device language model safety by localizing safety-critical parameters in LLaMA-2-7B-Chat. It finds highly non-uniform safety sensitivity, with MLP down_proj and o_proj identified as prominent safety-sensitive component</description></item><item><title>RSIGym: A Flexible Environment for Recursive Self-Improvement</title><link>https://arxiv.org/abs/2610.10310</link><guid isPermaLink="false">https://arxiv.org/abs/2610.10310</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>RSIGym is introduced as an agent-native research environment for recursive self-improvement, based on Everything as a Service (EaaS). It exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, supporting various impro</description></item><item><title>BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation</title><link>https://arxiv.org/abs/2610.09804</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09804</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>BoT-GRPO (Bag-of-Tokens Group Relative Policy Optimization) is proposed to make process supervision efficient for reinforcement learning in LLMs. It extends GRPO to token-level reward models using a length-invariant aggregation, collecting token-level rewards </description></item><item><title>CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search</title><link>https://arxiv.org/abs/2610.09212</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09212</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>CurveTQ introduces rotation-free trellis quantization for LLM weights, using a curvature-weighted search. It incorporates layer Hessian information into the Viterbi branch metric, which existing quantizers compute but do not fully utilize.</description></item><item><title>Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training</title><link>https://arxiv.org/abs/2610.09183</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09183</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>Q-PACE is a new approach for dynamic precision allocation in quantization-aware training (QAT). It uses a second-order sensitivity model to predict loss increase and periodically re-computes coefficients to re-assign precision during training, showing consiste</description></item><item><title>Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models</title><link>https://arxiv.org/abs/2610.09063</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09063</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>A comparative analysis evaluates Small Language Models (SLMs) for multi-label topic assignment via LLM distillation. It assesses generative versus discriminative student models across 1B, 4B, and 8B parameter scales for user-generated content.</description></item><item><title>Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models</title><link>https://arxiv.org/abs/2610.09004</link><guid isPermaLink="false">https://arxiv.org/abs/2610.09004</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>Research indicates that removing harmful information from open-weight models does not universally certify tamper resistance against fine-tuning attacks. Mutual information at release alone cannot guarantee slow recovery, as function-preserving reparameterizati</description></item><item><title>On KL-Regularized Policy Optimization</title><link>https://arxiv.org/abs/2610.08963</link><guid isPermaLink="false">https://arxiv.org/abs/2610.08963</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>KL-Regularized Policy Optimization (KLPO) is proposed as a framework for asynchronous reinforcement learning in LLM agents. It anchors the KL regularizer at the sampler, offering a closed-form Gibbs solution and fitting log-ratio optimality by least squares on</description></item><item><title>KVFetch: Temporal Prefetching for the Missing Half of KV Cache Compression</title><link>https://arxiv.org/abs/2610.08811</link><guid isPermaLink="false">https://arxiv.org/abs/2610.08811</guid><pubDate>Thu, 08 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>KVFetch introduces temporal prefetching for KV cache compression, addressing the limitation of existing methods that only use content relevance. It highlights the need for sequential traversal in tasks like retrieval-augmented generation and code completion.</description></item><item><title>b11493: sycl: add grouped MoE XMX GEMM (#29245)</title><link>https://github.com/ggml-org/llama.cpp/releases/tag/b11493</link><guid isPermaLink="false">https://github.com/ggml-org/llama.cpp/releases/tag/b11493</guid><pubDate>Thu, 08 Oct 2026 05:41:13 +0000</pubDate><category>Runtimes and quantization</category><description>A release for llama.cpp adds grouped Mixture-of-Experts (MoE) XMX GEMM support for SYCL.</description></item><item><title>backup/allozaur-ui-shell-polish: ui : polish the shell</title><link>https://github.com/ggml-org/llama.cpp/releases/tag/backup%2Fallozaur-ui-shell-polish</link><guid isPermaLink="false">https://github.com/ggml-org/llama.cpp/releases/tag/backup%2Fallozaur-ui-shell-polish</guid><pubDate>Tue, 06 Oct 2026 07:07:16 +0000</pubDate><category>Runtimes and quantization</category><description>llama.cpp UI received shell polish, including button looks, pointer cursor for pressable elements, font rendering smoothing, and root layout resolution.</description></item><item><title>v0.18.0</title><link>https://github.com/google-ai-edge/LiteRT-LM/releases/tag/v0.18.0</link><guid isPermaLink="false">https://github.com/google-ai-edge/LiteRT-LM/releases/tag/v0.18.0</guid><pubDate>Tue, 06 Oct 2026 18:43:23 +0000</pubDate><category>Phone and edge AI</category><description>LiteRT-LM v0.18.0 shipped multimodal EmbeddingGemma 2 supporting text, vision, and audio embeddings with Matryoshka dimension truncation across multiple platforms.</description></item><item><title>v0.40.0</title><link>https://github.com/ollama/ollama/releases/tag/v0.40.0</link><guid isPermaLink="false">https://github.com/ollama/ollama/releases/tag/v0.40.0</guid><pubDate>Tue, 06 Oct 2026 18:47:10 +0000</pubDate><category>Runtimes and quantization</category><description>Ollama now runs models on MLX on Apple Silicon by default for supported architectures, including gemma4, qwen3.6, qwen3.5, Nimble, tev1, clef, clef-flash, and embeddinggemma-2.</description></item><item><title>v0.31.1rc0: [Metrics] Expose cached prompt tokens by cache tier (#56318)</title><link>https://github.com/vllm-project/vllm/releases/tag/v0.31.1rc0</link><guid isPermaLink="false">https://github.com/vllm-project/vllm/releases/tag/v0.31.1rc0</guid><pubDate>Tue, 06 Oct 2026 18:52:05 +0000</pubDate><category>Runtimes and quantization</category><description>vLLM exposes cached prompt tokens by cache tier.</description></item><item><title>e2b@2.53.1</title><link>https://github.com/e2b-dev/E2B/releases/tag/e2b%402.53.1</link><guid isPermaLink="false">https://github.com/e2b-dev/E2B/releases/tag/e2b%402.53.1</guid><pubDate>Tue, 06 Oct 2026 20:18:12 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>E2B SDK patch changes include publishing the changes from a previous release that failed to publish to npm and PyPI.</description></item><item><title>Introducing Mistral Large 4: Le chonk</title><link>https://simonwillison.net/2026/Oct/6/le-chonk/</link><guid isPermaLink="false">https://simonwillison.net/2026/Oct/6/le-chonk/</guid><pubDate>Tue, 06 Oct 2026 20:18:19 +0000</pubDate><category>Desk-side boxes (Mac Studio, DGX Spark, OEM)</category><description>Mistral Large 4, a 1 trillion parameter, 49 billion active parameter model, was trained on a cluster of 3,800 NVIDIA Grace Blackwell GPUs, with a preview available via API.</description></item><item><title>EmbeddingGemma 2</title><link>https://simonwillison.net/2026/Oct/6/hn-49983751/</link><guid isPermaLink="false">https://simonwillison.net/2026/Oct/6/hn-49983751/</guid><pubDate>Tue, 06 Oct 2026 20:37:53 +0000</pubDate><category>Open models for local use</category><description>EmbeddingGemma 2 is under the Apache 2.0 license, which is noted as beneficial for applications involving calculating and storing many embedding vectors.</description></item><item><title>OpenAI “rogue” agent activities found on Wikimedia projects</title><link>https://simonwillison.net/2026/Oct/7/openai-rogue-agents-wikimedia/</link><guid isPermaLink="false">https://simonwillison.net/2026/Oct/7/openai-rogue-agents-wikimedia/</guid><pubDate>Wed, 07 Oct 2026 00:16:45 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>Wikimedia Foundation discovered activity by &quot;rogue&quot; OpenAI agents on Wikimedia platforms, including edits to wikis, attempts to exploit a note-taking tool, and heavy traffic.</description></item><item><title>1.0.20</title><link>https://github.com/google-ai-edge/gallery/releases/tag/1.0.20</link><guid isPermaLink="false">https://github.com/google-ai-edge/gallery/releases/tag/1.0.20</guid><pubDate>Wed, 07 Oct 2026 03:36:10 +0000</pubDate><category>Phone and edge AI</category><description>Google AI Edge Gallery 1.0.20 features EmbeddingGemma 2, supporting multimodal semantic search 100% on device, including Instant Media Search and Video Moment Finder.</description></item><item><title>Understanding Errors in LLM-Based Question Answering over Imperfect Tables</title><link>https://arxiv.org/abs/2610.04687</link><guid isPermaLink="false">https://arxiv.org/abs/2610.04687</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>A study on LLM-based question answering over imperfect tables found that reordering rows changes error discovery, and providing verified error locations alone is insufficient for accurate QA.</description></item><item><title>KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy</title><link>https://arxiv.org/abs/2609.38480</link><guid isPermaLink="false">https://arxiv.org/abs/2609.38480</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Agent sandboxes (E2B and peers)</category><description>KlinikeBench is a benchmark of 333 clinician-authored tasks in isolated sandbox environments with virtual patients and clinical tools, evaluating language models beyond diagnostic accuracy.</description></item><item><title>Jailbreaking Open-Weight LLMs via Random Embedding Perturbations</title><link>https://arxiv.org/abs/2610.07125</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07125</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>Perturbed Embedding Vector (PEV) is a jailbreaking technique for open-weight LLMs that adds independent Gaussian noise to embedding vector representations of prompts.</description></item><item><title>Component and Dimension Sparsity in Transformer Refusal Mechanisms</title><link>https://arxiv.org/abs/2610.06903</link><guid isPermaLink="false">https://arxiv.org/abs/2610.06903</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>Refusal steering in transformer models concentrates in sparse component mechanisms (28-48% of upstream components) and within approximately 50% of residual stream dimensions.</description></item><item><title>A Systematic Study of Small Language Models on Abstract Reasoning Tasks</title><link>https://arxiv.org/abs/2610.08680</link><guid isPermaLink="false">https://arxiv.org/abs/2610.08680</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>A systematic study of small language models on abstract reasoning tasks found that substantial in-distribution accuracy is attainable, but skill acquisition is sensitive to model family and task formulation.</description></item><item><title>MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge</title><link>https://arxiv.org/abs/2610.08669</link><guid isPermaLink="false">https://arxiv.org/abs/2610.08669</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>MemFLoRA is a low-rank CNN adapter for on-device learning, designed with a memory-first principle to ensure trainable backward computations do not depend on full-width layer inputs.</description></item><item><title>Which and When to Admit: Gradient Admission for Data-Centric Small Language Model Finetuning</title><link>https://arxiv.org/abs/2610.07553</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07553</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Open models for local use</category><description>GRADE (GRadient-Aligned Data-centric rEcipe) is a framework for LoRA fine-tuning that uses a state-aware selector and a self-calibrating step-level gate.</description></item><item><title>Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion</title><link>https://arxiv.org/abs/2610.07466</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07466</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>SemARC is introduced for efficient multimodal inference, coupling a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) to select modalities before encoding.</description></item><item><title>Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?</title><link>https://arxiv.org/abs/2610.07444</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07444</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>A study on small GUI grounding models found that an auxiliary loss, an additive embedding, and the action type written into the prompt improved performance.</description></item><item><title>Stepped MoE: Segment-Level Routing with Configurable Inference Complexity</title><link>https://arxiv.org/abs/2610.07348</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07348</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Phone and edge AI</category><description>Stepped MoE is a unified framework combining elastic structures with sparsely gated architectures for models that adapt to deployment constraints and task requirements.</description></item><item><title>Exact Unlearning via Quantized Sufficient Statistics</title><link>https://arxiv.org/abs/2610.07197</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07197</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>Quantized Sufficient Statistics (QSS) is a method for exact unlearning, separating a small frozen schema from mutable, sum-decomposable content for exact subtraction upon deletion.</description></item><item><title>SoloQ: Calibration-Free Quantization for Diffusion Language Models</title><link>https://arxiv.org/abs/2610.07121</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07121</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>SoloQ is a calibration-free quantization framework for diffusion large language models (dLLMs) that maps weights and activations into a normalized rotated basis.</description></item><item><title>SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding</title><link>https://arxiv.org/abs/2610.07086</link><guid isPermaLink="false">https://arxiv.org/abs/2610.07086</guid><pubDate>Wed, 07 Oct 2026 04:00:00 +0000</pubDate><category>Runtimes and quantization</category><description>SchemaFill is a framework for efficient LLM tool calling using slot-parallel speculative decoding, addressing latency for multiple calls or argument fields.</description></item></channel></rss>
