Skip to content

Science1 publisher3 min readPublished

A Nature Perspective wires a robot's internal signals into the policy that picks its actions

Sungwoo Lee and colleagues argue that embodied agents should carry an internal environment and let signals from it shape learning and choice. The journal's own editors say the demonstration is still to come.

The Scientist · Science desk

Illustration accompanying A Nature Perspective wires a robot's internal signals into the policy that picks its actions

What happened

  • Sungwoo Lee and colleagues propose in a Nature Machine Intelligence Perspective that embodied AI systems be given an internal environment and interoceptive inputs, monitoring internal states and updating their own goals.
  • The framework formalises robotic self-monitoring as a closed feedback loop, in which conditions in the environment alter internal signals and those signals then shape what the agent does.
  • Cybernetics was largely ignored as AI took shape after the 1956 Dartmouth workshop, which turned the field toward symbolic methods, knowledge representation and logical problem solving.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • decision A robotics team now has a design choice to argue about: battery and thermal telemetry as an interrupt sitting outside the controller, or as an input the policy conditions on. Those two builds fail in different ways.
  • constraint With the demonstration still outstanding, nobody can price interoception against the health monitor a robot already carries, so the framework competes on plausibility and on how cleanly it can be implemented.
  • contradiction The same editorial that describes cybernetics being sidelined also says control theory and robotics kept using it, so the novelty on offer is confined to learned policies.
  • capability If slow-moving internal variables do work as reference signals, an agent could adjust to an environment that changes under it without a retraining cycle.

The proposal turns on a distinction between a reward and a context variable. Earlier reinforcement learning work already used internal states as a source of reward for agents with open-ended goals [5]. Lee and colleagues keep that and add a second role: internal states as relatively stable contextual variables that influence learning and decision-making [6]. Their biological warrant is hunger and thirst, which in animals do not only generate reward but change what information is salient and which behaviours get priority [7].

In an agent, that would mean internal signals tuning policy selection, exploration strategy and memory updating [8]. Because those variables move slowly, the authors suggest they can act as a reference signal in environments that keep changing [9].

Robots already watch themselves. System health monitoring is long-standing practice, and in space exploration and disaster response it means tracking signals such as battery status, temperature and strain [4]. The framework formalises that monitoring as a closed feedback loop, with environmental conditions altering internal signals and those signals shaping actions in turn [3]. Under the proposal the same reading enters the policy as context for choosing what to do next [8].

Nature Machine Intelligence's editors are plain about the status of all this. "Future work will need to demonstrate whether the proposed framework can be of practical use in advancing robotics and physical AI, but the ideas are thought-provoking," the editorial said [10]. The comparison an engineer would want here is a robot run twice on one task, with its internal state visible to the policy in one arm and withheld in the other; the editorial does not report such a test and says the demonstration is future work [10].

The editorial's own history also complicates the idea that AI threw cybernetics away. Norbert Wiener set out the approach in the 1940s, with feedback loops at the centre of both living intelligence and machines [11]. After the 1956 Dartmouth workshop the field went to symbolic AI, knowledge representation, reasoning and logical problem solving, and embodiment and self-regulation got comparatively less attention [12]. Cybernetics was at least seven years old by then [14]. It did not vanish: robotics, control theory and autonomous systems carried it on [13]. What is being proposed is an import of those loops into learned policies.

On large language models, the objection recorded here is about environments rather than size. Such models typically pursue narrowly defined tasks inside well-specified computational settings, and extending them to reliable operation in physical environments remains challenging [15]. Real deployments bring changing conditions, limited resources and constraints on operation [16]. The case for interoception rests on those constraints, and it is structural: where a signal lives, and how it feeds back into action [3]. In organisms the same logic covers body temperature, oxygen levels and blood glucose, and keeping them in healthy ranges while the environment shifts is what motivates behaviour [17].

What to watch

  • A paired robot experiment from Lee's group or another lab: one task, one platform, internal state visible to the policy in one arm and withheld in the other.
  • Whether embodied benchmarks add episodes with depleting batteries and thermal limits, since full-charge fixed-length episodes cannot score an interoceptive agent at all.
  • Whether the framework yields a published formalism with measurable internal-state variables that a second lab can reimplement.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories