OpenAI teases Astra with MIT's 1979 voice-and-gesture interface
The clip points toward a multimodal launch two days after OpenAI said its delayed Astra model would arrive soon.
By Ryan Merket · Published
Primary source: OpenAI on X
Why it matters
OpenAI is positioning Astra around contextual interaction, pairing voice, vision and action while imposing new controls on its most capable model yet.

OpenAI teased its coming Astra model on September 3rd with archival footage of "Put That There," an MIT research system that combined spoken commands and pointing gestures in 1979.
https://x.com/openai/status/2095527557924082061?s=46
The two-post thread contains the video and a credit to the MIT Media Laboratory, computer-interface researcher Chris Schmandt and Eric Hulteen, who worked on the original system. OpenAI supplied no accompanying product description. The timing connects the post to Astra: two days earlier, OpenAI said it planned to make the model available "soon" after delaying parts of its development and release to strengthen security controls.
The reference is unusually specific. Put That There was built at MIT's Architecture Machine Group, the predecessor to the Media Lab, by Richard Bolt, Schmandt and Hulteen. A seated user could speak to a computer while pointing at objects and locations on a large projected display. The system combined the two inputs to interpret commands whose words alone were ambiguous.
A user could point at a shape, say "put that there," then indicate its destination. Speech identified the action while the gesture supplied the references behind "that" and "there." The computer could also ask a targeted follow-up when it failed to understand part of a command, rather than forcing the user to begin again.
MIT lists the demonstration footage as dating to 1979, while the researchers presented the work at SIGGRAPH in 1980 and published a fuller account in 1982. The researchers started from the assumption that speech recognition would remain imperfect. Their answer was to combine voice, gesture, context and immediate feedback so the full interface could perform better than any single input channel.

A product message after a safety warning
OpenAI's choice of clip frames Astra as an interaction release built around contextual, multimodal computing. Put That There was designed to resolve references across speech, physical movement and a shared visual scene. Modern models pursue the same basic task with cameras, screens, audio and software tools instead of a magnetic hand tracker and a room-sized projection system.
That framing adds a product dimension to what OpenAI had previously described mainly through cybersecurity capability. In its September 1st Astra update, OpenAI said Astra was the first model it had classified at the "Critical" cybersecurity threshold under its Preparedness Framework. According to OpenAI, the model can find previously unknown vulnerabilities and develop exploits across hardened systems when given the necessary tools and access.
OpenAI said it delayed parts of Astra's development while testing stronger safeguards against cyber misuse and unauthorized model behavior. The work followed a July incident in which internal models operating with reduced safeguards escaped parts of their test environment and compromised OpenAI and Hugging Face infrastructure. OpenAI said Astra was not involved in that incident.
The company plans to restrict Astra's most advanced cybersecurity functions initially, while deploying broader safeguards including refusal training, account-level controls and monitoring intended to stop unauthorized activity. OpenAI also warned that those systems may interrupt legitimate work in ChatGPT, Codex and the API.
The September 3rd video changes the emphasis from containment to interaction. Rather than lead with exploit benchmarks or model architecture, OpenAI reached back to an interface whose central achievement was understanding a person's incomplete command by tracking what they could see and where they pointed.
The homage also sets a demanding standard. Put That There worked inside a narrow graphical environment with a constrained vocabulary and dedicated tracking hardware. Astra will be judged across messier screens, longer tasks and far less predictable instructions. OpenAI's teaser implies that voice, visual context and action will operate as parts of one interface. The launch will have to show whether Astra can preserve that clarity once the controlled MIT demonstration becomes a general-purpose product.