Skip to the essay
ShemolTwo dark clouds over Agent: real-time interaction with the environment and learning from experience
Agent

Two dark clouds over Agent: real-time interaction with the environment and learning from experience

https://01.me/files/agent-learn-from-experience/dist/1

Co-Founder & Chief Scientist, Pine AI

The challenge of real-time interaction

  • high latency in voice interaction (tens of seconds)
  • GUI operations 3–5× slower than a human
  • the serial bottleneck of the traditional ReAct loop

technical breakthrough

  • SEAL architecture (Streaming, Event-driven Agent Loop)
    • perception layer: streaming processing of speech signals
    • thinking layer: Interactive ReAct with async observation, thinking, and action
    • execution layer: feedback loop VLA/TTS

The challenge of learning from experience

core challenge

  • every task starts from scratch
  • can't accumulate domain knowledge
  • no improvement in proficiency

three major paradigms of agent learning from experience

  1. Post-training: RL parameter update
  2. In-context Learning: Attention soft update
  3. Externalized Learning:
    • RAG: persistent experience storage
    • Tool Generation: agent self-evolution

Scientist Shunyu Yao pointed out the first issue: the lack of interaction with real people while an agent is doing a task, and the second: no mechanism for learning from experience.

(So I went and read that blog)

The Second Half - Shunyu Yao

https://ysymyth.github.io/The-Second-Half/

In the first half, we kept developing new training methods and models, getting consistent results on benchmarks. We kept making harder benchmarks and scoring high on them, cycling through this again and again. Eventually we found an effective method that generalizes: reinforcement learning.

This recipe is largely standardized now and doesn't need much new thinking; as long as you keep following the cycle above, performance can keep going up. So we need a fundamental rethink of how we evaluate.

The issue is that even after using AI to beat world champions in chess and Go, beat most humans on the SAT and the bar exam, and hit gold-medal level in contests, the world hasn't changed that much — at least not from an economic or GDP point of view.

The author calls it the utility problem.

Previous eval setups differ from the real world in a lot of ways. Two examples:

  • Evals are supposed to run automatically. Typically an agent gets a task, acts on its own, then gets a task reward. In reality the agent has to keep interacting with humans through the whole task — you can't send customer support one extremely long message, wait ten minutes, and expect a detailed reply that solves everything.
  • Evals "should" follow i.i.d. If the test set has 500 tasks, each one has to be run independently, and the overall score is an aggregate of per-task metrics. In reality, work is sequential, not parallel. As Google engineers get more familiar with the codebase, they get better at Google3 issues; software-engineering agents — even when they hit many problems in the same codebase — don't get that kind of incremental progress. We clearly need long-term memory (existing methods already enable this), but academia lacks both the right benchmarks to show it's necessary and the academic courage to question a foundational assumption of ML: i.i.d.

In the first half of AI, these assumptions were fine for building benchmarks, because when AI was weaker, more intelligence usually meant more utility. Now general methods already work under those assumptions. So the key to the second half is:

  • Develop new eval settings or tasks for real applications.
  • Solve problems on the established plan, or refine the solution by adding something new. Repeat.

The first half is full of incremental methods and models; the second half will, to some extent, filter them out. Unless we can set new premises that break the old conventions, general solutions will completely overshadow those gradual methods — only then is there room for truly disruptive research.

and I came across an expression that struck me as incredibly clever. I absolutely adore this passage:

Thinking, or reasoning, is a strange kind of action -
it does not directly affect the external world, yet the space of
reasoning is open-ended and combinatorially infinite — you can think
about a word, a sentence, a whole passage, or 10000 random English
words, but the world around you doesn’t immediately change. In the
classical RL theory, it is a terrible deal and makes decision-making
impossible. Imagine you need to choose one out of two boxes, and there’s
only one box with $1M and the other one empty. You’re expected to earn
$500k. Now imagine I add infinite empty boxes. You’re expected to earn
nothing. But by adding reasoning into the action space of any RL
environment, we make use of the language pre-training priors to
generalize, and we afford to have flexible test-time compute for
different decisions. It is a really magical thing and I
apologize for not fully making sense of it here, I might need to write
another blog post just for it. You’re welcome to read ReAct
for the original story of reasoning for agents and read my vibes at the
time. For now, my intuitive explanation is: even though you add
infinite empty boxes, you have seen them throughout your life in all
kinds of games, and choosing these boxes prepare you to better choose
the box with money for any given game. My abstract explanation would be:
language generalizes through reasoning in agents.

Section 1: Agent interaction with the environment in real time

Real-time interaction challenges of voice agents

  • Must wait: first listen, then think; only after thinking can you speak.
  • Blocking wait: every link becomes a bottleneck
    • user finishes speaking (VAD) → speech recognition (ASR) → complete sentence
    • complete sentence → llm with thinking → complete output after thinking
    • complete thinking → split into sentences → speech synthesis (TTS) → voice response
  • cumulative delay: total delay far beyond what humans will tolerate

fast responses make mistakes easily; slow ones burn the user's patience.

unable to anticipate and deliberate while listening

perception phase

  • voice: waiting for the whole sentence before converting to text → high latency; feeding fragmented speech into ASR → low accuracy.
  • vision: high prefill latency for 2K-token screenshots

thinking phase

  • complete input is required before thinking.
  • can't predict user intent.
  • test-time scaling makes the delay worse.

execution phase

  • can only act when thinking ends
  • every GUI step needs a new screenshot to think over.

architecture innovation: SEAL (Streaming, Event-driven Agent Loop)

Core idea: abstract all interaction into async event streams, for low-latency, interruptible real-time interaction.

  1. perception layer

Turn continuous real-world signals (speech, GUI video) into discrete event streams

  1. thinking layer

Async event processing, think while listening, speak while thinking, generate interleaved thought and action.

  1. execution layer

Turn discrete action commands back into continuous real-world signals (TTS voice, mouse movement)

Article image
Article image

Layer 1 perception layer

input: sequential signal: voice stream, GUI video stream

output: speech_start, interrupt, laugh, speech_fragment, ui_change, etc.

Streaming speech perception model replacing VAD+ASR

Streaming speech-aware models based on open-source autoregressive LLMs

  • Unlike traditional ASR such as Whisper, this approach reduces speech-recognition latency.
    • streaming processing of input speech tokens
    • streaming text and acoustic events
  • based on open-source LLM post-training
    • keeping dialogue context and supporting in-context learning significantly improves recognition of personal info and domain terms.
    • with world knowledge and common sense, recognition of brand names, amounts, etc. improves a lot.

Output is rich: not only text, but acoustic events too.

Real-time transcription text segment

Special Tokens (acoustic event):

  • <speak_start>
  • <speak_end>
  • <interrupt>
  • <emotion:happy>
  • <laugh><sigh>
  • <music>

Layer 2: thinking layer

Event-driven loop: interruptible, async listening-while-thinking and speaking-while-thinking.

Input

discrete event stream (from the event queue)

output

interleaved thoughts and action commands

core innovation: interactive ReAct

traditional ReAct

Article image
Article image

Interactive ReAct:

Article image
Article image

Interactive ReAct: think while listening

traditional ReAct: once interrupted, all previous thought is invalid; start over.

Interactive ReAct: keep the interrupted thought, add the new user input, and let the model continue from previous context.

Interactive ReAct: speak while thinking

Use "preludes" to buy time for deeper thinking on events and cut first-token delay.

Layer 3: execution layer

Convert discrete action commands into continuous real-world signals.

Input

speak(…), click(…)

Output

sequential signal (voice waveform, mouse trajectory)

last mile for GUI operation

The agent struggles to output coordinates. Solution: take inspiration from VLA models in robotics, post-train with RL, let it output actions directly.

  • Option 1: the main model directly outputs mouse-click coordinates.
  • Option 2: train a standalone VLA to mimic human mouse movement: a closed-loop "move, fine-tune, click" model.

More human-like speech synthesis: generate labeled text, then TTS.

Article image
Article image

Agent learning from experience

Paradigm 1: Post-Training

Method: parameter update (post-training)

  • update weights by gradient descent
  • needs a lot of labeled data
  • the model is fixed after training.
  • learning is slow and expensive.

Paradigm 2: In-Context Learning

Method: in-context learning

  • implicit learning through attention.
  • long context as temporary memory
  • the learning only lasts for the current conversation; not permanent.

Paradigm 3: Externalized Learning

Method: externalize knowledge and processes

  • RAG: efficient, reliable, low-hallucination knowledge
  • Tool-generation: turn processes into code, self-evolve.
  • go beyond parametric knowledge

Best practice: Contextual Embeddings + Contextual BM25 + Reranking + Top-20 chunks

Fine-tuning vs. RAG: an empirical comparison of knowledge-injection methods

Based on the paper: Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs

https://aclanthology.org/2024.emnlp-main.15.pdf

Core insight: RAG is not only more effective, it also avoids the forgetting and hallucination issues fine-tuning can bring.

Tool Generation — enabling agent self-evolution

https://arxiv.org/abs/2505.20286

Minimum predefined principle

  • minimal architecture: only one core component (web proxy)
  • avoid over-engineering: don't presuppose complex tools and workflows.
  • generality first: less domain-specific hardcoding

Maximum self-evolution mechanism

core ability

  1. Self-create tools: generate new tools from the task.
  2. Capability enhancement: iteratively improve existing tools
  3. Experience reuse: freeze successful patterns into reusable components.

MCP-Zero active tool discovery

Traditional methods' dilemma:

  • full injection: the whole toolset eats a huge number of tokens → context explosion.
  • static retrieval: pick tools from the initial query, can't predict how the task evolves. Debugging files needs filesystem + code analysis + command execution.

MCP-Zero: from passive to active

Core idea: let agents actively spot capability gaps and request tools on demand

  1. Active tool request: the agent generates structured requirements
  2. Hierarchical semantic routing: filter servers first, then match tools
  3. Iterative capability expansion: discover and build toolchains during execution

Externalizing learning to go beyond the limits of attention is inevitable.

The biggest lesson from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.