Timeline milestone
GPT-4o unifies real-time text, vision and voice interaction
-
Model Launch
GPT-4o unifies real-time text, vision and voice interaction
GPT-4o processed and generated combinations of text, audio and images with lower latency than earlier stitched-together systems.
OpenAI The “o” stood for “omni,” reflecting a shift from text chat toward more natural multimodal interaction. Voice and vision made AI assistants more immediate, while also raising stronger questions about impersonation, privacy, emotional dependence and the safeguards needed for live media generation.Sources & references 1 source