GPT‑Live‑1 in the API

arittr 54 points 55 comments September 11, 2026
openai.com · View on Hacker News

Discussion Highlights (14 comments)

loloisi

and yet another showcase of making automated restaurant reservations. It truly is the purpose of AGI, and all software ever, really, to automate that experience. It baffles me that the labs can't come up with more exciting use cases for voice api.

mesmertech

hopefully this come in openrouter api cause I'm not signing up for a specific provider's specific api platform, and have yet another thing that can bill me.

muddi900

The worst use case to show this.

dbbk

I've been waiting for this! I'm learning Spanish so I built an app to teach me Spanish, but hyperfocused on scenarios in my life, for example "watching a Barça match in a Barcelona bar". It does FSRS flashcard training, and live conversation practice. I think education is a very underexplored area for these live conversation models. Yes you can just use ChatGPT Live but that's freeform and unstructured, doesn't have a curriculum or can present supporting visuals, etc. On a grand scale if you can give children their own personal individual tutor rather than relying on group teaching alone, there could be a huge jump in successful education outcomes.

thiago_fm

Damn, this was a terrible showcase of their voice API, I really love their product. It's the fucking best, a life changer. But... what a meh showcase. Is somebody from OpenAI hiring? I can show how I use it to learn German, among other very interesting usages. Also show proper excitement etc... I think also a lot of real users could do better. It feels like they aren't real users of their own products...

idiliv

All voices offered sound human. I'd prefer a robotic voice, to avoid over-anthropomorphizing the AI.

gunalx

Can't wait till a customer rep is impossible to get to because all support is outsourced to ai.

almogo

Even as someone really AI-forwards, there are just not enough selling points for me here. I almost never want to talk to an AI. I just don’t believe I’ll have a useful voice interaction. Maybe agents are here to fix that, but theres 30 years of really negative precedent from robot telephone bots to overcome, and I don’t think some new API is going to change that overnight

Lucasoato

I've given a try to the demo in the webpage. I've asked to tell me which of the first generation pokemon started with the letter C, requesting it to tell me their names in reverse. It got stuck. I don't know if they are having some troubles with their demo environment due to the volume of requests in this specific moment, but I'm very hesitant to put something like this in production if it fails with this trivial example.

ndom91

Been wanting to try this out on top of https://github.com/TristanBrotherton/voicepe-realtime with the HomeAssistant VoicePE Anyone else hack something together with HA / the Voice PE yet? Looks like it needs a second model to do function calling, which gpt-realtime-2.5 didn't, and the Voice PE XMOS chip's audio pipeline might not be a great fit for full duplex back and forth, like what gpt-live-1 now supports.

agentdev001

Reposting my comment from https://news.ycombinator.com/item?id=49646963 Congrats on the launch here. I've been messing with this over the last few hours- super super cool. I was excitedly awaiting this hitting the API, because ofc there wasn't a super high fidelity option for drop-in voice interface in front of a given harness. This is blowing me away so far! (Side q, is there a single place one can watch for updates on the API- that actually covers everything that changes? IIRC there have been a couple of additions that you've tweeted- but never hit the API changelog ;] )

kailpa1

I think that this could be useful for the case of learning something by teaching it to someone else, and this someone else being the AI. We all know that learning-by-teaching is a great way to see the gaps in your knowledge and check whether you can explain the topic simple enough for the "student" to understand it. But finding the "student" is the hard thing in this process. Replacing the student with this model, and maybe a better reasoning model behind, sounds like a good enough replacement of a real person, for this case.

embedding-shape

> Reasoning & tool calling delegation: GPT‑Live‑1 can delegate reasoning and tool calls to a backend text model like GPT‑6 Astra or a third-party model. I played around a bit with this in Codex when it became available but even when you have Fast mode + Light reasoning, the mere idea that it passes off actual work to background sessions even for "change this line here" makes it a really frustrating experience. You can say "Update config here" then wait 2 minutes then finally it comes back, and most of those two minutes was overhead of agent<>sub-agent communication and passing the work, instead of just, you know, do the thing. I'm eagerly awaiting for this to get ready though, because being able to use tools like Houdini, Unreal Engine and Blender over MCP with this fast voice mode makes for great video game development environment, where you can playtest the game and talk with Codex at the same time, asking it to update stuff on the fly, granted you've setup things correctly.

varispeed

Can it infer someone's accent and correct it or tone of the voice or whether someone is talking in a mocking way? If not, then it's seems not there yet.

Semantic search powered by Rivestack pgvector
6,278 stories · 57,251 chunks indexed