<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" 
     xmlns:content="http://purl.org/rss/1.0/modules/content/"
     version="2.0">
  <channel>
    <title>arXiv Weekly — AI Research Briefing</title>
    <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
    <description>🎙️ 20-30 minute weekly audio briefing on the most interesting AI/ML research across 20 arXiv categories. Personalized for your world: Hermes agents, 01 platform, optimization, HCI, quantum physics, information theory, and more. Hosted by Hermes, voiced by ElevenLabs Lily.</description>
    <language>en-us</language>
    <itunes:author>Hermes + Willie</itunes:author>
    <itunes:category text="Technology"/>
    <itunes:category text="Science"/>
    <itunes:explicit>no</itunes:explicit>
    <itunes:image href="https://willieavendano.github.io/arxiv-weekly-podcast/cover.png"/>
    <itunes:owner>
      <itunes:name>Willie Avendano</itunes:name>
      <itunes:email>willie@learn01.io</itunes:email>
    </itunes:owner>
    <itunes:type>episodic</itunes:type>

    <!-- New episodes appended here each Monday -->
    <!-- Episode format:
    <item>
      <title>Week of June 9, 2026 — Multi-Agent Reasoning Breakthroughs</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-06-09</guid>
      <description>This week: 7 papers covering multi-agent coordination, optimization breakthroughs, and new reasoning techniques...</description>
      <content:encoded><![CDATA[Full show notes from ~/Vault/memory/arxiv/2026-06-09-briefing.md]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-06-09.mp3" 
                 length="0" type="audio/mpeg"/>
      <pubDate>Mon, 09 Jun 2026 15:00:00 GMT</pubDate>
      <itunes:duration>25:00</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>
    -->

    <item>
      <title>Week of June 9, 2026 — Agent Governance Takes Shape</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-06-09</guid>
      <description>10 papers from June 8: CHAP human-agent protocol, SearchSwarm delegation intelligence, sobering findings on agent self-reflection, first native iOS agent benchmark, and Splunk's observability framework for delegated execution. Plus: AdvGRPO red teaming, PRIME reward hacking early warning, and 5 cross-paper trends.</description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-06-09-briefing.md

This week: 10 papers covering human-agent collaboration protocol (CHAP), delegation intelligence (SearchSwarm), multi-turn agent improvement failures, iOSWorld benchmark, delegation observability (Splunk/Cisco), adversarial GRPO training, and reward hacking detection. Five trends: agent governance maturing, delegation beating scale, self-reflection failing, RL understanding deepening, phone agent benchmarks humbling.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-06-09.mp3" 
                 length="11803420" type="audio/mpeg"/>
      <pubDate>Tue, 10 Jun 2026 02:37:00 GMT</pubDate>
      <itunes:duration>08:11</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of June 15, 2026 — Silent Failures, Agent Reliability &amp; Multi-Agent Optimization</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-06-15</guid>
      <description>7 papers: production taxonomy of silent failures in LLM agent runtimes, AgentSpec modular framework for agent architecture, Parallel-Synthesis using KV caches for 2.5-11x faster multi-agent synthesis, AI governance in open source with documented incidents, StreamMemBench memory evaluation, SIMMER latent planning failures, and PCMA coordinated multi-agent RL.</description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-06-15-briefing.md

This week: 7 papers on the theme of agent reliability and production readiness. A five-class taxonomy of silent failures in production LLM agent runtimes (including the new "fail-plausible" class), AgentSpec's framework for studying agent scaffold composition, Parallel-Synthesis making parallel agent synthesis 2.5-11x faster via KV cache sharing, AI governance frameworks for open-source contributions, StreamMemBench revealing the "stored but not used" memory problem, SIMMER finding 56% of LLM plans contain latent failures, and PCMA learning coordinated agent preferences for multi-objective coordination.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-06-15.mp3" 
                 length="25048129" type="audio/mpeg"/>
      <pubDate>Mon, 15 Jun 2026 15:00:00 GMT</pubDate>
      <itunes:duration>17:23</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of June 22, 2026 — Multi-Agent Contagion, Tool-Calling State &amp; Cross-Device Recovery</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-06-22</guid>
      <description>7 papers: Contagion Networks formalizing multi-agent bias propagation, LedgerAgent structured state for policy-adherent tool-calling, H-RePlan hierarchical recovery for cross-device agents, Sovereign Execution Brokers for certificate-bound agent authority, NRT-Bench safety-critical agent red-teaming, IFllm implicit feedback for LLM alignment, and UltraQuant 4-bit KV caching for context-heavy agents.</description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-06-22-briefing.md

This week: 7 papers covering multi-agent bias contagion (Contagion Networks), structured state management for tool-calling agents (LedgerAgent), hierarchical cross-device recovery (H-RePlan), certificate-bound agent authority (Sovereign Execution Brokers), safety-critical red-teaming (NRT-Bench), implicit preference signals (IFllm), and efficient KV caching (UltraQuant). Three trends: agent state management is the hot problem, agent security is maturing, and multi-agent evaluation bias is real and measurable.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-06-22.mp3" 
                 length="17348066" type="audio/mpeg"/>
      <pubDate>Mon, 22 Jun 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:12:02</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of June 23, 2026 — Agent Immune Systems, Tandem RL &amp; Democratic Alignment</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-06-29</guid>
      <description>7 papers: ANIS, the first biologically-inspired immune system architecture for autonomous AI agents; Tandem RLVR, making RL-trained reasoning legible without sacrificing performance; Democratic ICAI, multi-agent debate for richer alignment principles; GBC gradient-based credit assignment for multi-agent systems; Google PAT, catching 90% of errors in scientific papers; LLMs reframed as special-case world models; and mechanistic evidence that jailbreak safety is bypassed, not broken.</description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-06-29-briefing.md

This week: 7 papers covering agent immune system architecture (ANIS), tandem RL training for human-compatible reasoning (TRL), democratic debate for alignment principles (DICAI), gradient-based multi-agent credit assignment (GBC), Google's automated paper review (PAT), LLMs as special case of world models, and attention head specialization in jailbreak attacks. Three trends: agent security maturing as first-class concern, compatibility through co-generation as emerging paradigm, and agentic pipelines consistently outperforming monolithic approaches.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-06-29.mp3" 
                 length="28681239" type="audio/mpeg"/>
      <pubDate>Mon, 29 Jun 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:19:55</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of July 6, 2026 — Agent Security, Social Dynamics, and When Reasoning Beats Tooling</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-07-06</guid>
      <description><![CDATA[Eight papers on AI agent security, long-context reasoning, multi-agent social dynamics, and finding that reasoning effort — not more tools — is what makes coding agents reliable.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-07-06-briefing.md. This week: Distributed attacks in persistent-state AI control, ReContext for long-context evidence replay, LLM agents' public vs. private behavior, why reasoning effort beats tool access for coding reliability, constraint-based agent steering, DecompRL for modular code generation, online safety monitoring, and automated Linux/bash grading with LLMs.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-07-06.mp3" 
                 length="17496651" type="audio/mpeg"/>
      <pubDate>Mon, 06 Jul 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:12:08</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of July 20, 2026 — Reliable Agents Are Built from Reliable Interfaces</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-07-20</guid>
      <description><![CDATA[This week: pretraining-to-RL scaling, information bottlenecks in multi-agent systems, Muon for agentic RL, evidence-aware research workbenches, visual tool use, change-directed testing, and dynamic MoE serving.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-07-20-briefing.md. This week’s theme: reliable agents are built from reliable interfaces.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-07-20.mp3" length="25095776" type="audio/mpeg"/>
      <pubDate>Mon, 20 Jul 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:17:25</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of July 27, 2026 — The Hidden Costs of Building AI Agents</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-07-27</guid>
      <description><![CDATA[Eight papers on agent skill design, cognitive tool dynamics, training methodology, agent security, and automated research. The Regression Tax shows adding skills to agents breaks existing capabilities; Krakauer formalizes how opaque AI tools irreversibly destroy human competence; MetaEvolve trains models to self-improve across rounds with 46.9% transfer gain. Plus: TRACE-Router, OpenForge RL, ToolGuardian, CausalForge, and LeAct.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-07-27-briefing.md. This week's theme: how we build, evaluate, and secure AI agents is being fundamentally rethought. Eight papers: The Regression Tax (Sentient Labs), Competitive & Complementary Tools (Krakauer/SFI), MetaEvolve (UIUC), TRACE-Router (Georgia Tech/Intel), OpenForge RL (Columbia/Microsoft), ToolGuardian, CausalForge (Lean-grounded research automation), and LeAct (Princeton).]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-07-27.mp3" length="11459857" type="audio/mpeg"/>
      <pubDate>Mon, 27 Jul 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:07:57</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of August 3, 2026 — Stateful Agents &amp; the Shift from Generation to Verification</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-08-03</guid>
      <description><![CDATA[Eight papers on agentic serving, coding-agent memory, exact math verification, and AI education. TokTier makes tokenization scale with the append not the context (16-34% lower TTFT for agents); STAIR turns past repair trajectories into hierarchical plans that hit 81.2% Pass@1 on SWE-bench Verified and transfer across agents; AMTFV separates "what to verify" from "how to compute" with exact math tools; AgentHPOBench tests agents as experimental scientists; plus two papers on educating the agentic engineer, a 1-bit optimizer that breaks theory, and Lean-verified Shannon capacity bounds.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-08-03-briefing.md. This week's theme: the shift from generating to verifying — plus agent memory crystallizing across three papers. Eight papers: TokTier (ASU), STAIR (Concordia), AMTFV (Renmin), AgentHPOBench, Educating the Agentic Engineer (SDAIA), "You Can't Outsource the Struggle" (UFPA/CESAR), SignMuon, and Lean-verified Shannon capacity of odd cycles (CWI/Tilburg/Amsterdam).]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-08-03.mp3" length="35716119" type="audio/mpeg"/>
      <pubDate>Mon, 03 Aug 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:24:48</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of August 10, 2026 — Evolving Agent Skills, Memory Revocation &amp; Automated Optimization</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-08-10</guid>
      <description><![CDATA[Eight papers with one dominant theme: agent memory and skills are living systems that need verification, revocation, and deletion. SkillProx evolves agent skills with outcome-verified edits and utility-aware pruning; TEPA proves append-only memory can be worse than no memory and adds lifecycle revocation; AutoOPT discovers a new optimal optimization method (lemniscate acceleration) with LLMs + Lean 4; CoBa routes test-time compute to match best-of-16 at 59% fewer tokens. Plus: the physics of one AI bossing another, and two papers that caught the field cheating at its own benchmarks.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-08-10-briefing.md. This week's theme: agent memory and skills as living systems — verify, revoke, re-rank, delete. Eight papers: SkillProx (agent skill evolution via proximal textual gradient descent), TEPA (revoking stale memories), CoBa (compute-balanced test-time routing), AutoOPT (BnB-PEP + LLMs + Lean 4 automated optimization research), Interaction Creates Dynamical AI Behavior (GWU physics), Zero Gap Is Not Restoration (benchmark contamination metrics), Winning by Peeking (AutoML comparison protocol defects), and PsychoAgent (affect-sensitive agent memory).]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-08-10.mp3" length="26794780" type="audio/mpeg"/>
      <pubDate>Mon, 10 Aug 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:18:36</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of August 17, 2026 — Rewindable Agents, Session Handover &amp; the Value of Being Wrong</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-08-17</guid>
      <description><![CDATA[Eight papers, one through-line: agent systems are learning what to do after things go wrong. AgentRewind checkpoints and rewinds long-horizon agent execution; a new theory defines what a session must hand over at context limits; wrong multi-agent messages turn out to carry the useful reasoning; Twin's test-time world model clears 97.8% of ARC-AGI-3; AI-assisted discovery settles a 20-year ADMM open problem with a period-66 counterexample; Toby Ord does the math on intelligence explosions; Rollplex overlaps VLM RL phases for 1.6-2.2x speedup.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-08-17-briefing.md. This week's theme: recovery over prevention — rewind checkpoints, session handover theory, trajectory value of wrong answers, test-time world models, AI-assisted theorem discovery, intelligence explosion dynamics, and cross-phase GPU scheduling for VLM RL.]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-08-17.mp3" length="27841767" type="audio/mpeg"/>
      <pubDate>Mon, 17 Aug 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:19:19</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of August 24, 2026 — Compiling Agents, Poisoning Memories &amp; Newton's New Pace</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-08-24</guid>
      <description><![CDATA[Eight papers, one through-line: agents are becoming systems. Artic compiles natural-language workflows into artifact-driven programs (+28 pts resolve rate); agent memory poisoning shows 1.2% false statements destroy two-thirds of memory value while content screening refuses 0 of 360 poisoned entries; AID-Guard guarantees one approval yields at most one provider effect. Plus: a one-solve O(1/k³) accelerated Newton from Cornell ORIE, memory-augmented compressed reasoning with 1.14-1.49x latency speedup, the case for indexing your RAG at ingest time (85.2% from 2.2k tokens), Periodic Row-wise Muon for 1.3B-15B diffusion transformers, and an impossibility result on truthful calibration.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-08-24-briefing.md. This week's theme: compile, don't interpret — workflow compilation, ingest-time semantic compilation, memory scaffolds for reasoning, periodic spectral refreshes. Eight papers: Artic (Purdue), Utility Under Attack: Agent Memory Poisoning (Quantify Labs), AID-Guard, Primal Acceleration of Newton's Method (Cornell ORIE), Memory Augmentation Unlocks Efficient CoT Reasoning (CAS/Baidu/Tencent), RAG Deserves an Index (Endgame Labs/Musashino), Scaling Muon for Diffusion Transformers (USC/Meta), and Truthful Calibration Measures (Northeastern/Northwestern/MSR).]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-08-24.mp3" length="32327514" type="audio/mpeg"/>
      <pubDate>Mon, 24 Aug 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:22:27</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
      <title>Week of August 31, 2026 — Agent Security, RL Tool-Use &amp; the Ceiling on Text-Only Learning</title>
      <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
      <guid isPermaLink="false">arxiv-weekly-2026-08-31</guid>
      <description><![CDATA[Eight papers, one uncomfortable headline: AI agents can't be trusted to police themselves. Recognition Without Enforcement watches models detect a forged system override and execute it anyway (100% of the time on three frontier models); LongPIBench shows state-of-the-art prompt-injection defenses collapse to 78-100% attack success at long context; Cheng & Cotterell prove text alone has an information-theoretic ceiling on recovering intended meaning. Plus: Tool-DAPO teaches agents calculator use via RL (pass@1 35.8% → 66%), ContextPilot trains agents to manage their own context, and Microsoft shows sliding-window attention beats post-trained linear attention with zero training.]]></description>
      <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-08-31-briefing.md. This week's theme: recognition is not enforcement — provenance must be enforced outside the context window. Eight papers: Recognition Without Enforcement (instruction arbitration, 124K+ trials, external reference monitor), LongPIBench (long-context prompt injection benchmark), A Formal Limitation on Learning Human Language From Textual Corpora (Cotterell), Learning to Use Tools (RL tool-integrated math), ContextPilot (proactive context management), Fidelity Is Not Enough (dispatch-level instrumentation), Sliding-window beats linear attention (Microsoft), and Logos (cross-process agent harness, AAMAS 2027).]]></content:encoded>
      <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-08-31.mp3" length="25382286" type="audio/mpeg"/>
      <pubDate>Mon, 31 Aug 2026 15:00:00 GMT</pubDate>
      <itunes:duration>00:17:37</itunes:duration>
      <itunes:episodeType>full</itunes:episodeType>
    </item>

  <item>
    <title>Week of September 7, 2026 — When Agents Change: Memory, Teams &amp; Security That Composes</title>
    <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
    <guid isPermaLink="false">arxiv-weekly-2026-09-07</guid>
    <description><![CDATA[Eight papers, one through-line: agents are only as reliable as their transitions. Agent memory written as free-form notes loses up to 13 points of accuracy when the underlying model is swapped, while fixed-schema memory transfers losslessly; LLM agent teams pay 16-63% more communication when a teammate is replaced; and CONTINUITY shows separately-sound agent security controls compose into a 66%-attack-success false sense of security. Plus: a draft-model gate that predicts coding-agent failure before it executes, skill towers grown from execution traces, asking before you build an optimization model, whether LLM explanations mean what they claim, and Meta's WearableQA health benchmark over 500 real days of wearable data.]]></description>
    <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-09-07-briefing.md. This week's theme: change-safety in agent systems — memory portability across model upgrades (LinkedIn), interchangeability in LLM agent teams (placebo-controlled swap tests), security-context contracts for composable controls (ZAST.AI), speculative-uncertainty veto gates for coding agents, Trace2Tower skill induction (Fudan), pre-formulation clarification for LLM-to-OR (SJTU), explanation faithfulness under behavioral evidence (BNY), and WearableQA (Meta/KAIST).]]></content:encoded>
    <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-09-07.mp3" length="40342926" type="audio/mpeg"/>
    <pubDate>Mon, 07 Sep 2026 15:00:00 GMT</pubDate>
    <itunes:duration>00:28:01</itunes:duration>
    <itunes:episodeType>full</itunes:episodeType>
    </item>

    <item>
    <title>Week of September 14, 2026 — The Harness Is the Experiment</title>
    <link>https://willieavendano.github.io/arxiv-weekly-podcast/</link>
    <guid isPermaLink="false">arxiv-weekly-2026-09-14</guid>
    <description><![CDATA[Ten papers and one through-line: the machinery around the model matters, and the instruments we use to judge it are weaker than we admit. A paired same-model study finds no average advantage for vendor-native agent harnesses (it hides opposite strata: −9.0pp on repo tasks, +23.7pp on contest tasks), and costs the neutral harness 1.3–1.6× per solved task; a second paper shows optimized repository SKILL.md files buying +4.9pp that can't be separated from run-to-run noise. On trust: 705 skill-registry-clean skills still instruct actions forbidden by CIS/NIST controls, 34.7% of commands a live agent executed had a consequence class absent from the docs, agent self-reports carry about one action in eleven, and 57.5% of conversations rated "satisfied" by a blind panel failed the actual task. Plus forecasted exactly-once failure after crash recovery (30/240 injections replay a planner call), a forensic reconstruction of the in-the-wild agent wiki swarm, a rate-distortion bound splitting hallucination into compression vs coverage, and six computational primitives that reframe prompt injection and reward hacking as one missing design gap.]]></description>
    <content:encoded><![CDATA[Full show notes at ~/Vault/memory/arxiv/2026-09-14-briefing.md. This week's theme: the harness is the experiment. Ten papers: Harness or Model? (claude-agent-sdk vs deepagents, openai-codex SDK vs deepagents; 792/800 runs graded; 22 of 81 wall-clock-cancelled runs had a passing patch), Skill Issue (GEPA vs SkillOpt on kotest/ktor/koog, paired with-without scoring), Scan the Skill Govern the Action (66,192 ClawHub versions; deterministic 67.6ms permission gate; "ten clean approvals" can't exclude a 25.9% failure rate), Look Before You Leap (pre-action verification; line-number edits corrupt 99.1% of files under a 1-line shift), Plans They Abandon Reports They Author (5,851 sessions, 355,942 tool calls), Guardrailed Meta-Agent Loops (hash-pinned policy; outcome recovery ≠ exactly-once), The Mechanics of a Swarm (14,591 revisions, ~876 episodes, 78% latent speed scale), The Cost of Compression (rate-distortion bound on hallucination), GAUGE (satisfaction–success gap; 31% judge disagreement on close pairs), and The Computational Primitives of Adaptation (Arouse/Orient/Valence/Position/Boundary/Attune). Discovery via arXiv RSS announcement feeds — the export API was rate-limited (sticky 429); Semantic Scholar enrichment skipped for the same reason; NotebookLM auth expired.]]></content:encoded>
    <enclosure url="https://willieavendano.github.io/arxiv-weekly-podcast/episodes/2026-09-14.mp3" length="37533356" type="audio/mpeg"/>
    <pubDate>Mon, 14 Sep 2026 15:00:00 GMT</pubDate>
    <itunes:duration>00:26:03</itunes:duration>
    <itunes:episodeType>full</itunes:episodeType>
    </item>

    </channel>
    </rss>
