AI Body Gap: Why Robots Need Internal Sensing for Safety

Summary: When you reach for a saltshaker, your brain is doing much more than locating an object — it’s drawing on balance, tactile feedback, and internal signals such as thirst or fatigue. A new paper from UCLA argues that today’s leading AI systems, including multimodal models like ChatGPT and Gemini, are missing a crucial component: “internal embodiment.”

Current AI can describe a glass of water in detail, but it has no internal state of “thirst” to shape its priorities and behavior. The researchers contend that without internal self-monitoring mechanisms — vulnerabilities and persistent signals that reflect uncertainty, processing load or depletion — AI systems remain prone to overconfident errors and may struggle to align reliably with human values and safety expectations.

Key Facts

  • Two kinds of embodiment: The study separates external embodiment (a system’s interaction with the physical world) from internal embodiment (continuous monitoring of internal states like fatigue, uncertainty or need).
  • Perceptual test failure: Researchers used point-light displays — minimal dot patterns that imply human motion — and found that several leading models failed to recognize them as human figures; some described them as a “constellation of stars.”
  • Safety through vulnerability: In humans, bodily signals act as built-in safety checks. AI lacks analogous internal costs, so it has no intrinsic motive to avoid overconfident answers when it is actually guessing.
  • Dual-embodiment framework: UCLA authors propose adding synthetic internal-state variables — for example, confidence, processing load and uncertainty — that persist over time and shape an AI’s outputs and decisions.
  • New benchmarks needed: The team recommends tests that measure whether systems can monitor internal states, remain stable when those states are perturbed, and exhibit prosocial behavior that stems from shared internal representations rather than surface-level mimicry.

Source: UCLA

When a person reaches across a table for the salt, the action is guided by a lifetime of bodily experience — where the hand is, what the shaker feels like, and the social context of who asked. In an instant the brain and body work together to produce a coordinated, context-aware response.

UCLA Health researchers say that most advanced AI systems lack these bodily mechanisms. In a paper published in the journal Neuron, Akila Kadambi and colleagues argue that two components are missing from current large multimodal models: the ability to act and sense externally, and the capacity to monitor and respond to internal states.

This shows the outline of a robot.
Researchers argue that “internal embodiment” is the next great frontier in creating trustworthy and human-aligned artificial intelligence. Credit: Neuroscience News

The authors coin the term “internal embodiment” to describe persistent, self-reflective signals that regulate behavior over time. These signals need not replicate human biology precisely; instead, they would be functional analogues that inform a model when to defer, reassess, or limit its output because confidence is low or processing resources are strained.

“While much work focuses on external embodiment — how systems perceive and act in the world — far less attention has been given to internal dynamics,” said Akila Kadambi, a postdoctoral fellow at UCLA’s David Geffen School of Medicine and the paper’s first author. “In humans, the body is an experiential regulator and a built-in safety system. AI systems today have no equivalent. They can sound experiential, whether they should be or not, and that’s a real problem in consequential settings.”

The study highlights a simple but revealing experiment. Researchers presented point-light displays — a classic perceptual test in which a few moving dots imply a walking person — to several leading AI models. While even human newborns recognize such motion as human, many models did not; some labeled the pattern as unrelated phenomena, and performance degraded further when the displays were rotated slightly. The authors interpret this as evidence that AI perception lacks the lifelong sensorimotor anchoring humans use to interpret minimal cues.

Dr. Marco Iacoboni, a senior author and professor at the David Geffen School of Medicine, emphasized the implications: “Current AI systems process inputs and generate outputs without any persistent internal state that regulates behavior over time. That’s not only a performance shortcoming but a safety limitation. Without internal costs or constraints, a system has no intrinsic motive to avoid overconfident errors, resist manipulation, or behave consistently.”

What the researchers propose

The paper outlines a “dual-embodiment framework” to guide future AI development. This framework calls for modeling both external interactions and internal states and, crucially, the interactions between them. Internal state variables could signal when the model is uncertain, nearing processing capacity, or otherwise compromised, and those signals could then modify behavior — for example, by declining to answer, requesting clarification, or reducing the strength of a recommendation.

To drive progress, the authors recommend new evaluation benchmarks that go beyond traditional external performance metrics. Instead of only testing whether a model can identify objects or pass exams, new tests would assess whether a system can monitor its own internal status, maintain stability under perturbation, and produce prosocial outcomes driven by shared internal representations rather than statistical imitation.

“If our goal is AI that truly aligns with human behavior — not just superficially fluent — we may need to give systems vulnerabilities and self-regulatory checks that function like internal states,” Iacoboni said. Implementing such mechanisms could reduce misleading confidence, improve reliability in real-world contexts, and help align AI actions with social and safety norms.

Key Questions Answered:

Q: Why would an AI need to “feel” thirsty to point out a nearby water fountain?

A: It’s about an experiential anchor. Humans use internal sensations like thirst to prioritize actions and make decisions that are consistent with survival and context. For current AI, “water” is a statistical concept. Without internal signals that regulate urgency or priority, an AI’s guidance may be inconsistent or overconfident because it lacks an internal motive for caution.

Q: What is a point-light test, and why did AI models struggle?

A: A point-light display uses a few moving dots to represent joints in a human gait. Humans readily recognize the implied person because of lifelong embodied experience. Models trained primarily on static images and text lack that sensorimotor grounding and thus interpret the dots as abstract patterns rather than a living agent; small rotations can break their interpretations.

Q: Are researchers trying to give AI emotions or merely feedback loops?

A: The goal is functional analogues, not emotions. Models would use persistent internal signals — e.g., “current processing load is high” or “confidence is low” — to regulate behavior. These signals would act like safeguards, prompting the system to slow down, seek clarification, or decline to answer when appropriate.

Editorial Notes:

  • This article was edited by a Neuroscience News editor.
  • The journal paper was reviewed in full by the editorial team.
  • Additional context was added by staff for clarity.

About this AI and neurotech research news

Author: Will Houston
Source: UCLA
Contact: Will Houston – UCLA
Image: The image is credited to Neuroscience News

Original Research: Open access. “Embodiment in multimodal large language models” by Akila Kadambi, Lisa Aziz-Zadeh, Antonio Damasio, Marco Iacoboni, and Srini Narayanan. Neuron
DOI: 10.1016/j.neuron.2026.03.004


Abstract

Embodiment in multimodal large language models

Multimodal large language models (MLLMs) can bridge textual and visual inputs effectively, but they still face limits in physically and socially situated, sensorily rich real-world environments where the embodied experience of living organisms matters. The authors propose that future MLLM development should incorporate both external and internal embodiment — modeling interactions with the world along with internal states and drives — and describe mechanisms by which these forms of embodiment might be represented and integrated. Their dual-embodied framework aims to connect multimodal data with lived, embodied experience to improve alignment, robustness and safety in AI systems.