OpenAI has released two new speech-to-text models, with prices reduced by 25%.
According to TrendForce Monitoring, OpenAI has released two speech-to-text models, GPT Transcribe and GPT Live Transcribe. The former handles audio files and batch tasks, while the latter is used for live captions in scenarios like live broadcasts and phone calls.
Both models can incorporate audio topics, keywords, and language cues, focusing on improving recognition accuracy for short sentences, numbers, technical terms, multiple accents, and noisy environments.
In a study by Artificial Analysis, GPT Transcribe has a word error rate of 3.31%, which is 0.7 percentage points lower than the previous generation GPT-4o Transcribe. The price has also been reduced by 25%, charging $4.5 per 1000 minutes of audio.