OpenAI unveiled GPT-4o, the latest iteration of its artificial intelligence (AI) that powers ChatGPT. It is being made available gradually to all users, including those using the free version.

This is the first model capable of processing images and voices in real time. In earlier versions, ChatGPT relied on other AI models to interpret voice commands and images. This change is expected to make ChatGPT even faster, with several different features.

One of the most anticipated features was the new ChatGPT's ability to "read" websites in real time. This function already existed in other AI tools on the market, but it never worked as expected for very long in those AIs. However, other developments that most of the public did not expect arrived and caught many people by surprise.

New: Ability to Interact in Real Time

With the arrival of GPT-4o, ChatGPT gains the ability to interact in real time, including audio and image features that allow it to play audio and interpret photos and videos, such as those on YouTube, during conversations. It also enables impressively accurate simultaneous translation between two people who do not speak the same language, as shown in the video below.

Live demo of GPT-4o realtime translation
GPT-4o demonstration video from OpenAI's official channel.

New: Image Recognition

Improvements in image recognition were highlighted during the presentation. One practical example was identifying and solving a simple mathematical equation, with ChatGPT assisting in solving a problem involving analytic geometry.

Math problems with GPT-4o
GPT-4o demonstration video from OpenAI's official channel.

New: Recognition of External Three-Dimensional Objects

The GPT-4o model recognizes external three-dimensional objects by processing visual information in real time. Using computer vision techniques, the model can analyze images or videos, identify three-dimensional objects in the environment, and describe them in any language.

Point and Learn Spanish with GPT-4o
GPT-4o demonstration video from OpenAI's official channel.

New: Interpretation of Facial Expressions

Through advanced image-processing and machine-learning techniques, the new model can identify facial patterns and interpret the emotions expressed on a person's face. In addition, GPT-4o can identify and interpret different situations, as in the example below, where the AI gives a young man tips on how to conduct himself in a job interview and what he should wear.

Interview Prep with GPT-4o
GPT-4o demonstration video from OpenAI's official channel.

OpenAI is beginning the rollout of GPT-4o's text and image features this Monday. Users of the free version will be able to access it, although the message limit was not specified, while ChatGPT Plus subscribers will have a higher limit. The use of GPT-4o with voice commands will become available to ChatGPT Plus subscribers in the coming weeks.

The launch of the new product comes amid OpenAI's efforts to stay ahead of growing competition in the race for AI leadership. Rival companies such as Google and Meta have dedicated themselves to developing increasingly robust language models that can be applied across a range of AI products.

The latest GPT release could represent an advantage for Microsoft, which invested significantly in OpenAI with the goal of integrating its AI technology into its own products.