https://01.me/files/agent-learn-from-experience/dist/1
Co-Founder & Chief Scientist, Pine AI
The challenge of real-time interaction
- high latency in voice interaction (tens of seconds)
- GUI operations 3–5× slower than a human
- the serial bottleneck of the traditional ReAct loop
technical breakthrough
- SEAL architecture (Streaming, Event-driven Agent Loop)
- perception layer: streaming processing of speech signals
- thinking layer: Interactive ReAct with async observation, thinking, and action
- execution layer: feedback loop VLA/TTS
The challenge of learning from experience
core challenge
- every task starts from scratch
- can't accumulate domain knowledge
- no improvement in proficiency
three major paradigms of agent learning from experience
- Post-training: RL parameter update
- In-context Learning: Attention soft update
- Externalized Learning:
- RAG: persistent experience storage
- Tool Generation: agent self-evolution
Scientist Shunyu Yao pointed out the first issue: the lack of interaction with real people while an agent is doing a task, and the second: no mechanism for learning from experience.
(So I went and read that blog)
The Second Half - Shunyu Yao
https://ysymyth.github.io/The-Second-Half/
In the first half, we kept developing new training methods and models, getting consistent results on benchmarks. We kept making harder benchmarks and scoring high on them, cycling through this again and again. Eventually we found an effective method that generalizes: reinforcement learning.
This recipe is largely standardized now and doesn't need much new thinking; as long as you keep following the cycle above, performance can keep going up. So we need a fundamental rethink of how we evaluate.
The issue is that even after using AI to beat world champions in chess and Go, beat most humans on the SAT and the bar exam, and hit gold-medal level in contests, the world hasn't changed that much — at least not from an economic or GDP point of view.
The author calls it the utility problem.
Previous eval setups differ from the real world in a lot of ways. Two examples:
- Evals are supposed to run automatically. Typically an agent gets a task, acts on its own, then gets a task reward. In reality the agent has to keep interacting with humans through the whole task — you can't send customer support one extremely long message, wait ten minutes, and expect a detailed reply that solves everything.
- Evals "should" follow i.i.d. If the test set has 500 tasks, each one has to be run independently, and the overall score is an aggregate of per-task metrics. In reality, work is sequential, not parallel. As Google engineers get more familiar with the codebase, they get better at Google3 issues; software-engineering agents — even when they hit many problems in the same codebase — don't get that kind of incremental progress. We clearly need long-term memory (existing methods already enable this), but academia lacks both the right benchmarks to show it's necessary and the academic courage to question a foundational assumption of ML: i.i.d.
In the first half of AI, these assumptions were fine for building benchmarks, because when AI was weaker, more intelligence usually meant more utility. Now general methods already work under those assumptions. So the key to the second half is:
- Develop new eval settings or tasks for real applications.
- Solve problems on the established plan, or refine the solution by adding something new. Repeat.
The first half is full of incremental methods and models; the second half will, to some extent, filter them out. Unless we can set new premises that break the old conventions, general solutions will completely overshadow those gradual methods — only then is there room for truly disruptive research.
and I came across an expression that struck me as incredibly clever. I absolutely adore this passage:
Thinking, or reasoning, is a strange kind of action -
it does not directly affect the external world, yet the space of
reasoning is open-ended and combinatorially infinite — you can think
about a word, a sentence, a whole passage, or 10000 random English
words, but the world around you doesn’t immediately change. In the
classical RL theory, it is a terrible deal and makes decision-making
impossible. Imagine you need to choose one out of two boxes, and there’s
only one box with $1M and the other one empty. You’re expected to earn
$500k. Now imagine I add infinite empty boxes. You’re expected to earn
nothing. But by adding reasoning into the action space of any RL
environment, we make use of the language pre-training priors to
generalize, and we afford to have flexible test-time compute for
different decisions. It is a really magical thing and I
apologize for not fully making sense of it here, I might need to write
another blog post just for it. You’re welcome to read ReAct
for the original story of reasoning for agents and read my vibes at the
time. For now, my intuitive explanation is: even though you add
infinite empty boxes, you have seen them throughout your life in all
kinds of games, and choosing these boxes prepare you to better choose
the box with money for any given game. My abstract explanation would be:
language generalizes through reasoning in agents.
Section 1: Agent interaction with the environment in real time
Real-time interaction challenges of voice agents
- Must wait: first listen, then think; only after thinking can you speak.
- Blocking wait: every link becomes a bottleneck
- user finishes speaking (VAD) → speech recognition (ASR) → complete sentence
- complete sentence → llm with thinking → complete output after thinking
- complete thinking → split into sentences → speech synthesis (TTS) → voice response
- cumulative delay: total delay far beyond what humans will tolerate
fast responses make mistakes easily; slow ones burn the user's patience.
unable to anticipate and deliberate while listening
perception phase
- voice: waiting for the whole sentence before converting to text → high latency; feeding fragmented speech into ASR → low accuracy.
- vision: high prefill latency for 2K-token screenshots
thinking phase
- complete input is required before thinking.
- can't predict user intent.
- test-time scaling makes the delay worse.
execution phase
- can only act when thinking ends
- every GUI step needs a new screenshot to think over.
architecture innovation: SEAL (Streaming, Event-driven Agent Loop)
Core idea: abstract all interaction into async event streams, for low-latency, interruptible real-time interaction.
- perception layer
Turn continuous real-world signals (speech, GUI video) into discrete event streams
- thinking layer
Async event processing, think while listening, speak while thinking, generate interleaved thought and action.
- execution layer
Turn discrete action commands back into continuous real-world signals (TTS voice, mouse movement)

Layer 1 perception layer
input: sequential signal: voice stream, GUI video stream
output: speech_start, interrupt, laugh, speech_fragment, ui_change, etc.
Streaming speech perception model replacing VAD+ASR
Streaming speech-aware models based on open-source autoregressive LLMs
- Unlike traditional ASR such as Whisper, this approach reduces speech-recognition latency.
- streaming processing of input speech tokens
- streaming text and acoustic events
- based on open-source LLM post-training
- keeping dialogue context and supporting in-context learning significantly improves recognition of personal info and domain terms.
- with world knowledge and common sense, recognition of brand names, amounts, etc. improves a lot.
Output is rich: not only text, but acoustic events too.
Real-time transcription text segment
Special Tokens (acoustic event):
<speak_start><speak_end><interrupt><emotion:happy><laugh><sigh><music>
Layer 2: thinking layer
Event-driven loop: interruptible, async listening-while-thinking and speaking-while-thinking.
Input
discrete event stream (from the event queue)
output
interleaved thoughts and action commands
core innovation: interactive ReAct
traditional ReAct

Interactive ReAct:

Interactive ReAct: think while listening
traditional ReAct: once interrupted, all previous thought is invalid; start over.
Interactive ReAct: keep the interrupted thought, add the new user input, and let the model continue from previous context.
Interactive ReAct: speak while thinking
Use "preludes" to buy time for deeper thinking on events and cut first-token delay.
Layer 3: execution layer
Convert discrete action commands into continuous real-world signals.
Input
speak(…), click(…)
Output
sequential signal (voice waveform, mouse trajectory)
last mile for GUI operation
The agent struggles to output coordinates. Solution: take inspiration from VLA models in robotics, post-train with RL, let it output actions directly.
- Option 1: the main model directly outputs mouse-click coordinates.
- Option 2: train a standalone VLA to mimic human mouse movement: a closed-loop "move, fine-tune, click" model.
More human-like speech synthesis: generate labeled text, then TTS.

Agent learning from experience
Paradigm 1: Post-Training
Method: parameter update (post-training)
- update weights by gradient descent
- needs a lot of labeled data
- the model is fixed after training.
- learning is slow and expensive.
Paradigm 2: In-Context Learning
Method: in-context learning
- implicit learning through attention.
- long context as temporary memory
- the learning only lasts for the current conversation; not permanent.
Paradigm 3: Externalized Learning
Method: externalize knowledge and processes
- RAG: efficient, reliable, low-hallucination knowledge
- Tool-generation: turn processes into code, self-evolve.
- go beyond parametric knowledge
Best practice: Contextual Embeddings + Contextual BM25 + Reranking + Top-20 chunks
Fine-tuning vs. RAG: an empirical comparison of knowledge-injection methods
Based on the paper: Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
https://aclanthology.org/2024.emnlp-main.15.pdf
Core insight: RAG is not only more effective, it also avoids the forgetting and hallucination issues fine-tuning can bring.
Tool Generation — enabling agent self-evolution
https://arxiv.org/abs/2505.20286
Minimum predefined principle
- minimal architecture: only one core component (web proxy)
- avoid over-engineering: don't presuppose complex tools and workflows.
- generality first: less domain-specific hardcoding
Maximum self-evolution mechanism
core ability
- Self-create tools: generate new tools from the task.
- Capability enhancement: iteratively improve existing tools
- Experience reuse: freeze successful patterns into reusable components.
MCP-Zero active tool discovery
Traditional methods' dilemma:
- full injection: the whole toolset eats a huge number of tokens → context explosion.
- static retrieval: pick tools from the initial query, can't predict how the task evolves. Debugging files needs filesystem + code analysis + command execution.
MCP-Zero: from passive to active
Core idea: let agents actively spot capability gaps and request tools on demand
- Active tool request: the agent generates structured requirements
- Hierarchical semantic routing: filter servers first, then match tools
- Iterative capability expansion: discover and build toolchains during execution
Externalizing learning to go beyond the limits of attention is inevitable.
The biggest lesson from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.