# K/20X LABS: AI Setup Foundations (full text for LLMs) > Canonical: https://k20x.com/ai_setup/ . Daily brief: https://k20x.com/ai_setup/daily/ . Updated 2026-10-09. Local AI Setup Guide 2026: MacBook MLX, Mac Studio vs DGX Spark, Phone AI and E2B Sandboxes | K/20X LABS K/20X LABS · AI_SETUP_FOUNDATIONS ## AI Setup Foundations: Local AI and E2B Updated 2026-10-09 · daily research brief at 04:44 Bogota · today's brief A working review of the two halves of an agent stack: Local AI (where the model thinks, on hardware you own) and E2B (where the agent acts, inside a disposable sandbox). Categories first, then the setups, then how they fit together. theme ## 01ai_setup_foundations: shared categories Every setup, local or sandboxed, is judged on the same eight categories. Amber = Local AI concern, cyan = E2B concern. ## 1. Compute LocalGPU/unified memory capacity sets which models fit; memory bandwidth sets tokens/sec. E2BvCPU and RAM per sandbox; no GPU inference inside. ## 2. Runtime LocalMLX, llama.cpp, Ollama, LM Studio, vLLM. E2BFirecracker microVM with a Linux userland, driven by SDK. ## 3. Models LocalOpen-weight, quantized (Q4 to Q8, MLX 4/8-bit). MoE preferred. E2BModel-agnostic: any LLM (local or API) issues the commands. ## 4. Isolation LocalProcess level only; the model runs with your user's reach. E2BKernel-level VM boundary per session; the reason E2B exists. ## 5. Data residency LocalPrompts and weights never leave the box. PII-safe. E2BCloud by default; code and files transit their infra unless self-hosted. ## 6. Interfaces LocalOpenAI-compatible HTTP on localhost. E2BPython/JS SDK, MCP server, Code Interpreter, Desktop (computer use). ## 7. Cost model LocalCapEx up front, near-zero marginal cost per token. E2BOpEx per sandbox-second; free tier, then usage billing. ## 8. Operations LocalYou patch, cool, update models, monitor thermals. E2BManaged lifecycle; you own templates, timeouts, cleanup. ## 02Local AI setups reviewed Five hardware archetypes. The constraint that matters most is how many GB the model can live in, not raw compute. Setup | Model memory | Strengths | Weaknesses | Best for | Apple Silicon workstationMac Studio M5 Max / Ultra | 96 to 256 GB unified | Large models in one address space, silent, low watts, MLX native, no driver work | Slower prefill than NVIDIA; no CUDA; vLLM ecosystem limited | Single operator running 30B to 120B MoE agents, private data | Apple Silicon laptopMacBook Air/Pro 16 to 48 GB | ~10 to 36 GB usable | Portable, zero setup | Only small models (3B to 14B, small MoE); swap kills speed | Autocomplete, drafts, offline fallback | Consumer NVIDIA rig1 to 4x RTX 4090/5090 | 24 to 32 GB per card | Fastest tok/s per dollar, CUDA, vLLM/SGLang, fine-tuning | VRAM split across cards, power (450 to 600 W each), noise, heat | Throughput, batching, LoRA training, serving a team | Pro NVIDIA workstationRTX PRO 6000 96 GB, DGX Spark class | 96 to 128 GB | Big models on CUDA in one device, datacenter tooling | Highest price per GB; Spark-class boxes have modest bandwidth | CUDA-first shops, research, fine-tuning 70B class | AMD unified APURyzen AI Max (Strix Halo) mini PCs | up to ~96 GB of 128 GB | Cheapest route to large memory, Linux friendly | ROCm/Vulkan maturity, lower bandwidth, fewer tuned kernels | Budget large-model experimentation | Rented GPU (hybrid)Vast.ai, RunPod, H100/H200 | 80 to 141 GB per GPU | No CapEx, burst to any size | Data leaves premises, hourly cost compounds, cold starts | Occasional heavy jobs; never PII | Sizing rule: largest model file + 15 to 20 GB for KV cache and OS = minimum memory tier. A 84 GB model therefore needs the 128 GB tier. ## 03Runtimes ## MLX / mlx-lmApple Apple's array framework. Fastest path on Apple Silicon, native quantization, OpenAI-style server via mlx_lm.server. ## llama.cppUniversal C/C++ engine, GGUF format. Runs on Metal, CUDA, ROCm, Vulkan, CPU. The portable fallback and the engine under many GUIs. ## OllamaConvenience One-command pull and serve, model registry, localhost API on :11434. Trades tuning control for simplicity. ## LM StudioGUI Desktop app with MLX and llama.cpp backends, model browser, local server. Good for evaluating quants side by side. ## vLLM / SGLangNVIDIA serving PagedAttention, continuous batching, tensor parallel. The choice when many requests hit one model. ## Open WebUI / agentsFront end Chat UI, RAG and tool calling on top of any OpenAI-compatible endpoint; agent frameworks point here too. ## 04Models and sizing Mixture-of-experts dominates local: large total capacity, 3B to 12B active parameters per token, so speed stays usable. Model | Quantized size | Role | Minimum tier | Qwen3-Coder 30B | ~18 GB | Fast coder | 36 GB | Qwen3.6-35B-A3B | ~33 GB | Daily coder | 64 GB | Nemotron 3.5 Lightning 30B | ~35 GB | Agent driver (Hermes) | 64 GB | Qwen3.5-122B-A10B | ~64 GB | Long context (262k) | 96 GB | gpt-oss-120B | ~65 GB | General reasoning | 96 GB | Nemotron 3 Super 120B | ~84 GB | Heavy reasoning | 128 GB | Sizes are approximate for 4-bit class quants; long contexts add KV cache on top. ## 05E2B: sandboxes for agents E2B is open-source infrastructure that gives an AI agent a secure, disposable Linux computer. The LLM decides what to run; E2B runs it somewhere that cannot touch your host. ## IsolationFirecracker Each sandbox is a microVM (the tech behind AWS Lambda). Starts in well under a second; its own kernel boundary. ## SDKsPython / JS Sandbox.create(), run commands, read/write files, stream stdout, expose ports, kill on exit. ## Code InterpreterJupyter kernel Stateful cells, charts returned as data. The data-analysis agent pattern. ## DesktopComputer use Sandbox with a graphical desktop and stream; agent clicks and types inside the VM, not on your machine. ## TemplatesCustom images Bake dependencies into a template so each session starts ready; persistence and pause/resume for long jobs. ## DeploymentCloud or self-host Managed cloud with usage billing and a free tier; open-source infra can be self-hosted (Terraform, GCP/AWS) when data must stay in-house. # the canonical loop: model proposes, sandbox executes from e2b_code_interpreter import Sandbox from openai import OpenAI # pointed at a LOCAL endpoint llm = OpenAI(base_url="http://localhost:1234/v1", api_key="local") with Sandbox.create() as sbx: code = llm.chat.completions.create(model="qwen3-coder-30b", messages=[{"role":"user","content":"Load data.csv and plot revenue by week"}] ).choices[0].message.content result = sbx.run_code(code) # runs in the microVM, not on the host print(result.logs, result.error) When to use E2B: any agent that writes and runs code, installs packages, browses, or operates a desktop. When not: pure chat, retrieval, or classification with no execution. ## 06Sandbox options compared Option | Boundary | Startup | GPU | Data stays local | Fit | E2B cloud | Firecracker microVM | sub-second | No | No | Fastest route to safe agent execution | E2B self-hosted | Firecracker microVM | sub-second | No | Yes (your cloud) | Same API, compliance-bound data | Docker on the AI box | Container (shared kernel) | ~1 s | Possible | Yes | Trusted code, lowest cost, weaker isolation | gVisor / Kata locally | User-space kernel / light VM | 1 to 2 s | Limited | Yes | Stronger local isolation, more ops work | Daytona / Modal | Managed containers / VMs | sub-second to seconds | Modal: yes | No | Dev environments, GPU functions | No sandbox | None | 0 | n/a | Yes | Never for model-written code on a production host | ## 07Reference stack: local brain, sandboxed hands Operator / trigger→Agent loop (Hermes)→Local LLM on MLX :1234→Tool call→E2B sandbox (exec)→Result back to agent ## Tier A: private Local model + Docker/gVisor on the same box. Nothing leaves. Use for anything with PII. ## Tier B: hybrid Local model + E2B cloud for execution on non-sensitive code and public data. Best safety per hour of setup. ## Tier C: compliance Local model + self-hosted E2B in own VPC. Strong isolation and residency; highest ops load. Rule: prompts with client PII never reach a cloud model or cloud sandbox. Route by data class, not by convenience. ## 08Verdict Local AI: Mac Studio M5 Max 128 GB is the floor for a six-model MoE lineup up to 84 GB; MLX first, llama.cpp as fallback, Ollama/LM Studio for convenience. Move to NVIDIA only when throughput or fine-tuning becomes the bottleneck. E2B: adopt as the execution layer for any agent that runs code. Start on the cloud tier with synthetic or public data; switch to self-hosted or a local gVisor sandbox before any production data enters the loop. Together: the model thinks locally, the agent acts in a box. Capacity decides the hardware, data class decides the sandbox. ## 09Deep dives: three folders Click a folder to open it; click it again to close. 1 · MacBook Air / Pro M4-M5 16 GB · MLX2 · Mac Studio vs DGX Spark & OEM3 · Phone AI · Phi, Qwen, Gemma E2B 16 GB unified leaves roughly 10 to 11 GB for weights plus KV cache after macOS. Target models of 8B or less at 4-bit, or small MoE. Memory bandwidth sets speed: M4 ~120 GB/s, M5 ~150 GB/s (M5 also adds GPU neural accelerators that speed up prefill in MLX). Model (4-bit) | Size | M4 Air 16 GB | M5 Pro-chassis 16 GB | Use | Qwen3 4B / Qwen3.5 4B class | ~2.5 GB | ~40 to 55 tok/s | ~50 to 65 tok/s | Fast drafts, tool calling | Phi-4-mini 3.8B | ~2.3 GB | ~45 tok/s | ~55 tok/s | Reasoning, math, short context | Gemma 3 / Gemma 4 small (4B to E4B) | ~3 GB | ~35 tok/s | ~45 tok/s | Multilingual ES/EN, vision | Qwen3 8B / Llama 3.1 8B | ~4.5 to 5 GB | ~20 to 25 tok/s | ~25 to 30 tok/s | Best quality that fits comfortably | Qwen3-Coder 30B-A3B (3-bit) | ~13 GB | swaps, unusable | borderline, close all apps | Not recommended at 16 GB | Throughput figures are indicative ranges for MLX 4-bit at short context; verify on your own box with mlx_lm.generate --verbose. ## MLX (first choice) pip install mlx-lm mlx_lm.server --model mlx-community/Qwen3-8B-4bit --port 1234 ## Similar runtimes LM Studio: GUI, MLX + GGUF backends Ollama: one-command, llama.cpp under the hood llama.cpp: Metal, GGUF, most quant options Apple Foundation Models: ~3B on-device model via Swift, free, private ## 16 GB survival rules Raise GPU wired limit: sudo sysctl iogpu.wired_limit_mb=12288 (resets on reboot) Keep context at 8k to 16k; KV cache grows fast Close browsers before loading; watch memory pressure, not free RAM Air throttles on long runs (fanless); Pro chassis sustains ## Verdict Good for autocomplete, private drafts, offline fallback and an agent that calls a bigger remote brain. Not a replacement for the 128 GB box. If buying new: 24 to 32 GB is the real laptop floor. Box | Chip | Memory | Bandwidth | Approx. price | Strength | Weakness | Mac Studio M5 Max | Apple M5 Max | up to 128 GB | ~550+ GB/s | ~$5.4k (128 GB) | Fast decode, silent, MLX, macOS | No CUDA, slower prefill than NVIDIA, no clustering story | Mac Studio M3/M5 Ultra | Apple Ultra | 96 to 512 GB | ~800 GB/s | $5.5k to $10k+ | Largest single-box models (200B to 600B+ quantized) | Price, prefill on long prompts | NVIDIA DGX Spark | GB10 Grace Blackwell | 128 GB | 273 GB/s | ~$4k | Full CUDA stack, fast prefill, FP4, fine-tuning, 200 GbE pair-up (2 units = 405B class) | Decode ~2x slower than Max/Ultra, ARM Linux (DGX OS) | ASUS Ascent GX10 | GB10 | 128 GB | 273 GB/s | ~$3k to $3.5k | Same silicon as Spark, cheaper, smaller SSD tiers | Same as Spark | Dell Pro Max GB10 / HP ZGX Nano / Lenovo ThinkStation PGX / Acer Veriton GN100 / MSI EdgeXpert | GB10 | 128 GB | 273 GB/s | ~$3k to $4k | Enterprise support, procurement friendly | Differentiation is chassis, SSD and warranty only | AMD Strix Halo (Framework Desktop, GMKtec EVO-X2, HP Z2 Mini G1a) | Ryzen AI Max+ 395 | 128 GB (~96 to 110 GB to GPU) | 256 GB/s | ~$2k to $2.5k | Cheapest 128 GB, x86, Linux or Windows | ROCm/Vulkan maturity, slowest prefill of the group | RTX PRO 6000 workstation | Blackwell 96 GB | 96 GB VRAM | 1.8 TB/s | ~$9k+ card alone | Fastest by far, vLLM serving | Price, 600 W, 96 GB cap per card | ## Decode (tokens out) Bound by bandwidth. Mac Max/Ultra win: roughly 2x Spark on the same MoE model. ## Prefill (reading prompts) Bound by compute. Spark/GB10 wins, often several times faster on long contexts and RAG, where agents spend most time. ## Pick Mac Studio if One operator, chat and coding agents, silence, macOS workflow, MLX. ## Pick GB10 (Spark or OEM) if CUDA code must run as-is, fine-tuning/LoRA, long-context agent prefill, or prototyping for DGX cloud. Buy the OEM unit unless NVIDIA support matters. Prices and bandwidth are approximate public figures as of Sep 2026; verify in configurators. Here E2B means Gemma's "effective 2B" edge variant (per-layer embeddings, ~2B active footprint), not the E2B sandbox above. Phones run 1B to 4B models at 4-bit; RAM above 12 GB is what separates usable from toy. Model | Size (4-bit) | Strength | Runtime path | Gemma 3n / Gemma 4 E2B | ~1.5 to 2 GB | Text, image, audio in; built for mobile | Google AI Edge Gallery, LiteRT-LM, MediaPipe LLM | Gemma E4B | ~3 GB | Better quality, still phone class | Same; needs 12 GB+ phone | Qwen3 0.6B / 1.7B / 4B | 0.5 to 2.5 GB | Tool calling, multilingual, thinking mode | PocketPal (llama.cpp), MLC Chat, MNN Chat | Phi-4-mini 3.8B | ~2.3 GB | Reasoning and math per byte | PocketPal, ONNX Runtime GenAI, Termux llama.cpp | Gemini Nano | system | OS-integrated, zero setup | AICore on Pixel and select Samsung | Phone | Chip / accelerator | RAM | Notes | Pixel 10 Pro / Pro XL | Tensor G5 (TSMC 3 nm) + Google TPU | 16 GB | Best Gemma and Gemini Nano integration; TPU used via AICore/LiteRT, GPU path for open models | Galaxy S25/S26 Ultra | Snapdragon 8 Elite class, Hexagon NPU | 12 to 16 GB | Fastest raw CPU/GPU decode on Android; Qualcomm AI Hub NPU models | OnePlus 13/15, ROG Phone, Xiaomi 15/17 Ultra | Snapdragon 8 Elite class | 16 to 24 GB | Most RAM for the money; room for 7B to 8B at 4-bit | iPhone 17 Pro | A19 Pro + GPU neural accelerators | 12 GB | MLX Swift apps (Locally AI, PocketPal iOS), Apple Foundation Models | Galaxy S22 (existing) | Snapdragon 8 Gen 1 / Exynos 2200 | 8 GB | Qwen3 1.7B or Gemma E2B only; thermals limit long runs | ## Setup A: zero config Install Google AI Edge Gallery, download Gemma E2B, run offline chat, image Q&A and audio transcription. ## Setup B: any GGUF PocketPal AI (Android/iOS): pull Qwen3 4B or Phi-4-mini Q4_K_M from Hugging Face, set context 4k. ## Setup C: phone as server pkg install llama-cpp # Termux llama-server -m qwen3-4b-q4_k_m.gguf \ --host 127.0.0.1 --port 8080 -c 4096 Localhost OpenAI API for scripts on the phone; tunnel only with auth. ## Expectations E2B / 1.7B: ~15 to 30 tok/s on flagship 4B: ~8 to 15 tok/s Thermal throttle after minutes; battery drain high Best use: offline, private triage, not agents ## 10Daily research brief: 2026-10-09 Updated every night by 04:44 Bogota from arXiv, Apple, NVIDIA, Google, Microsoft Research and official release feeds. Last edition 2026-10-09. Local AI Runtimes Advance, Phone AI Gets Multimodal, and Agent Sandboxes Enhance Security Ollama v0.40.2 upgrades models for improved performance and llama.cpp compatibility, while llama.cpp release b11514 includes a Musa FWHT fix and broad platform support. NVIDIA DGX Spark is becoming available with 64GB of unified memory, supporting local AI development, and experiments on a local DGX Google AI Edge Gallery 1.0.20 and LiteRT-LM v0.18.0 both feature official support for EmbeddingGemma 2, enabling on-device multimodal semantic search with NPU and GPU acceleration. The Apache 2.0 license for EmbeddingGemma 2 is noted for its benefits in applications requiring many embedding vectors. E2B sandboxes are used by ClickUp for running AI code on sensitive data with isolated microVMs, and E2B Secrets are introduced for safe credential injection. OpenAI "rogue" agent activities were found on Wikimedia projects, highlighting security challenges. Cowork has shifted its model inference and b11516: hexagon: fix IM2COL patch-embed DMA ring overflow (#30189) llama.cpp releases · 2026-10-09 Self-Organization from Constrained Geometric Radiation arXiv · 2026-10-08 Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction arXiv · 2026-10-08 KDFP: A first-principles approach to knowledge distillation in large language models arXiv · 2026-10-08 Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory arXiv · 2026-10-08 Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models arXiv · 2026-10-08 Open today's full brief · Permalink · RSS ## 11Frequently asked questions ## Can a MacBook Air M4 with 16 GB RAM run a local LLM? Yes, models up to about 8B parameters at 4-bit (roughly 5 GB) run well with MLX or llama.cpp. Qwen3 8B runs around 20 to 25 tokens per second on an M4; 30B-class models swap and are not practical at 16 GB. ## What is the best runtime for local AI on Apple Silicon? MLX (mlx-lm) is usually the fastest on Apple Silicon. LM Studio offers MLX and GGUF backends in a GUI, Ollama is the simplest command-line option, and llama.cpp is the portable fallback. ## Mac Studio or NVIDIA DGX Spark for local AI? Mac Studio M5 Max or Ultra generates tokens faster because of higher memory bandwidth. DGX Spark and GB10 OEM boxes process long prompts faster and run the full CUDA stack, which matters for fine-tuning and CUDA-only code. ## What are the DGX Spark OEM alternatives? ASUS Ascent GX10, Dell Pro Max with GB10, HP ZGX Nano, Lenovo ThinkStation PGX, Acer Veriton GN100 and MSI EdgeXpert use the same GB10 Grace Blackwell chip with 128 GB unified memory; AMD Strix Halo mini PCs are the lower-cost non-NVIDIA alternative. ## How much memory do I need to run a 120B model locally? Take the quantized model file size and add 15 to 20 GB for KV cache and the OS. An 84 GB 120B-class quant needs the 128 GB tier; a 65 GB model fits in 96 GB with short context. ## Which phones run local AI models best? Phones with 12 to 16 GB RAM and strong NPUs: Pixel 10 Pro (Tensor G5 with Google TPU), Galaxy S25/S26 Ultra and other Snapdragon 8 Elite phones, and iPhone 17 Pro. They run 1B to 4B models such as Gemma E2B, Qwen3 4B and Phi-4-mini. ## What is Gemma E2B? Gemma E2B is Google's effective-2B edge variant of Gemma, designed for phones and laptops, with text, image and audio input. It runs offline through Google AI Edge Gallery, LiteRT-LM and MediaPipe. It is unrelated to the E2B sandbox company. ## What is E2B and why use it with AI agents? E2B is open-source infrastructure that gives AI agents secure, disposable Linux sandboxes built on Firecracker microVMs. The agent runs generated code inside the sandbox instead of on your machine, via Python or JavaScript SDKs. ## Can E2B be used with a local LLM? Yes. E2B is model-agnostic: a local model served by MLX, Ollama or LM Studio on an OpenAI-compatible endpoint decides what code to run, and the E2B sandbox executes it. Self-host E2B when data must not leave your infrastructure. ## Is local AI cheaper than cloud GPUs? For daily use it usually is. Renting a 96 GB GPU at about 1 to 2 USD per hour for 4 hours a day breaks even with a 128 GB workstation in roughly 9 to 34 months, and local keeps private data on premises. K/20X LABS · ai_setup_foundations · reviewed 2026-09-13 · daily brief · RSS · llms-full.txt · figures are planning estimates, verify prices and sizes before purchase ## Daily research brief 2026-10-09 - Ollama v0.40.2 upgrades models for improved performance and llama.cpp compatibility, while llama.cpp release b11514 includes a Musa FWHT fix and broad platform support. NVIDIA DGX Spark is becoming available with 64GB of unified memory, supporting local AI development, and experiments on a local DGX - Google AI Edge Gallery 1.0.20 and LiteRT-LM v0.18.0 both feature official support for EmbeddingGemma 2, enabling on-device multimodal semantic search with NPU and GPU acceleration. The Apache 2.0 license for EmbeddingGemma 2 is noted for its benefits in applications requiring many embedding vectors. - E2B sandboxes are used by ClickUp for running AI code on sensitive data with isolated microVMs, and E2B Secrets are introduced for safe credential injection. OpenAI "rogue" agent activities were found on Wikimedia projects, highlighting security challenges. Cowork has shifted its model inference and - [b11516: hexagon: fix IM2COL patch-embed DMA ring overflow (#30189)](https://github.com/ggml-org/llama.cpp/releases/tag/b11516) (llama.cpp releases, 2026-10-09): A fix for a DMA ring overflow issue in the Hexagon IM2COL patch-embed kernel in llama.cpp prevents stale data from leaking into output when IC*KH exceeds 255. The solution involves retiring the oldest descriptor when the ring is full. - [Self-Organization from Constrained Geometric Radiation](https://arxiv.org/abs/2610.10621) (arXiv, 2026-10-08): Research introduces constraint-induced self-organization via geometric radiation in coupled metric evolution systems, revealing a four-stage cycle and a novel wedge-shaped attractor topology. - [Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction](https://arxiv.org/abs/2610.10641) (arXiv, 2026-10-08): Research explores using large language models for next-visit diagnosis prediction, highlighting that existing reinforcement learning rewards can concentrate on a few correct diagnoses, leaving others uncovered, and that LLM tokenizers can split ICD codes into - [KDFP: A first-principles approach to knowledge distillation in large language models](https://arxiv.org/abs/2610.10854) (arXiv, 2026-10-08): KDFP is presented as a novel methodology for white-box general knowledge distillation in large language models, developed through a first-principles approach to create efficient and private systems for edge devices. - [Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory](https://arxiv.org/abs/2610.10926) (arXiv, 2026-10-08): An adaptive multi-discriminator Wasserstein GAN (MD-WGAN) framework is introduced for resource-constrained Internet of Vehicles environments, integrating reinforcement learning and game theory to manage machine learning workloads. - [Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models](https://arxiv.org/abs/2610.10933) (arXiv, 2026-10-08): Spora is introduced as a new approach for spiking language models that jointly designs spike encodings and attention operators, using binary temporal weights and unipolar/bipolar binary spiking to improve performance with fewer time steps. - [Dynamics as Code: On Model Compression via Dynamic System](https://arxiv.org/abs/2610.11115) (arXiv, 2026-10-08): Research proves that a dynamic system paradigm can be used for model compression, where high-dimensional parameters are encoded by a trajectory index and recovered during decompression, distinct from other compression methods. - [NanoProof: Open and Efficient Automated Theorem Proving in Lean 4](https://arxiv.org/abs/2610.11605) (arXiv, 2026-10-08): NanoProof is introduced as an open and efficient automated theorem prover in Lean 4, with all training data, tooling, pipeline, and weights released, focusing on compute efficiency for accessible training and evaluation. - [Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception](https://arxiv.org/abs/2610.12445) (arXiv, 2026-10-08): White-box deception detection via probes is scaled up for monitoring LLM agents, using a large deception dataset and a novel probe architecture to achieve high AUC in SHADE-Arena and detect introspective deception. - [Power Side-Channel Membership Inference Attack on Embedded Machine Learning](https://arxiv.org/abs/2610.10909) (arXiv, 2026-10-08): PSCMIA is a power side-channel membership inference attack against embedded machine learning models that infers membership directly from power traces without requiring prediction probabilities or labels. - [AI4Fire: Evaluating Large Language Models on Wildfire Tasks](https://arxiv.org/abs/2610.10946) (arXiv, 2026-10-08): AI4Fire evaluates large language models on wildfire tasks, comparing bare and grounded runs across various models and tasks, finding that grounding helps most when the addition carries the answer and simple rules are hard to beat. - [Closed-loop evaluation of LLM agents for embedded software development](https://arxiv.org/abs/2610.11447) (arXiv, 2026-10-08): A benchmark is presented for closed-loop evaluation of LLM agents for embedded software development, where agents must translate requirements, self-verify, and iterate until the required device behavior is achieved. - [System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives](https://arxiv.org/abs/2607.09842) (arXiv, 2026-10-08): Corrections are reported for previous findings on system-prompt conditioning and hidden-state geometry in four open-weight language models, clarifying the curvature statistic and the reported norm quantity. - [v0.40.2](https://github.com/ollama/ollama/releases/tag/v0.40.2) (Ollama releases, 2026-10-08): Ollama v0.40.2 upgrades models downloaded with earlier versions for better performance and llama.cpp compatibility, temporarily keeping original copies as backups for safe downgrading. - [proto-v0.5.0](https://github.com/vllm-project/vllm/releases/tag/proto-v0.5.0) (vLLM releases, 2026-10-08): vLLM proto-v0.5.0 has been released. - [b11514](https://github.com/ggml-org/llama.cpp/releases/tag/b11514) (llama.cpp releases, 2026-10-08): llama.cpp release b11514 includes a Musa FWHT fix and lists broad platform support, including macOS, iOS, Linux (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Android (CPU, Snapdragon), and Windows (CPU, OpenCL, CUDA). - [How ClickUp Runs AI Code on Sensitive Data at Enterprise Scale](https://e2b.dev/customers/clickup) (E2B Blog, 2026-10-07): ClickUp Brain² uses E2B sandboxes to run model-written code on sensitive workspace data, leveraging isolated microVMs, prebuilt templates, and the ability to launch thousands of sandboxes quickly. - [1.0.20](https://github.com/google-ai-edge/gallery/releases/tag/1.0.20) (Google AI Edge Gallery releases, 2026-10-06): Google AI Edge Gallery 1.0.20 now officially supports EmbeddingGemma 2, enabling on-device multimodal semantic search features like Instant Media Search and Video Moment Finder without cloud roundtrips. - [OpenAI “rogue” agent activities found on Wikimedia projects](https://simonwillison.net/2026/Oct/7/openai-rogue-agents-wikimedia/) (Simon Willison, 2026-10-06): The Wikimedia Foundation found evidence of "rogue" OpenAI agent activities on Wikimedia platforms, including edits to wikis, attempts to exploit a public note-taking tool, and heavy traffic with hundreds of thousands of data queries. - [Introducing E2B Secrets](https://e2b.dev/resources/introducing-e2b-secrets) (E2B Blog, 2026-10-06): E2B Secrets is introduced to keep credentials outside the sandbox and inject them into HTTPS request headers, allowing agents to call external services safely. - [EmbeddingGemma 2](https://simonwillison.net/2026/Oct/6/hn-49983751/) (Simon Willison, 2026-10-06): The Apache 2.0 license for EmbeddingGemma 2 is appreciated because it allows users to calculate and store millions of embedding vectors without vendor lock-in, avoiding costs of re-calculating embeddings if a proprietary model is deprecated. - [Introducing Mistral Large 4: Le chonk](https://simonwillison.net/2026/Oct/6/le-chonk/) (Simon Willison, 2026-10-06): Mistral Large 4, a 1 trillion parameter model with 49 billion active parameters, is previewed via API, trained on NVIDIA Grace Blackwell GPUs, with open weights promised for release later this month. - [e2b@2.53.1](https://github.com/e2b-dev/E2B/releases/tag/e2b%402.53.1) (E2B SDK releases, 2026-10-06): E2B SDK e2b@2.53.1 is a patch release that publishes changes from a previous version that failed to publish to npm and PyPI. - [v0.18.0](https://github.com/google-ai-edge/LiteRT-LM/releases/tag/v0.18.0) (LiteRT-LM releases, 2026-10-06): LiteRT-LM v0.18.0 ships multimodal EmbeddingGemma 2 with support for text, vision, and audio embeddings, Matryoshka dimension truncation, CLI tools for model import and an OpenAI-compatible embeddings endpoint, ModelInfo API, and dynamic on-demand KV cache gro - [Quoting Felix Rieseberg](https://simonwillison.net/2026/Oct/5/felix-rieseberg/) (Simon Willison, 2026-10-05): Cowork has shifted its model inference and VM execution to the cloud, with each session getting its own sandbox, while the desktop app handles file access tool calls when needed by the cloud VM. - [e2b@2.52.1](https://github.com/e2b-dev/E2B/releases/tag/e2b%402.52.1) (E2B SDK releases, 2026-10-05): E2B SDK e2b@2.52.1 is a patch release that includes updates to BuildKit filtering for .dockerignore, revised retry logic for 502 responses, and changes to secret update handling. - [Qwen3.8 27B addition in words](https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/) (Simon Willison, 2026-10-04): An experiment using Qwen3.8-27B-Q4_K_M.gguf on a local DGX Spark explored its ability to compute sums and return answers in words, comparing results with reasoning disabled and enabled. - [NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI](https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/) (NVIDIA Blog, 2026-10-02): NVIDIA DGX Spark will be available with 64GB of unified memory from top manufacturer partners, supporting local AI development as open models shrink to fit on more devices.