Find a model

Search models and open their dedicated profile. Use the arrow keys to navigate and Enter to open.

Grok Voice STT 1.0

xAI · 1 configuration

Intelligence—
Input / 1M—
Output / 1M—
Context15K

Tariff: ZenMux · Source: models-dev · USD per million tokens

Thinking configuration

What does more thinking buy you?

Compare this model’s measured thinking configurations.

◌

No pair of measurements available

Change the filters or pick another benchmark. Missing values are not estimated.

Selected: thinking default · Click a point to switch configuration

Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.

This model is in the catalog but has no imported benchmarks. Capabilities are not estimated.

Specifications from the catalog

Inputaudio
Outputtext
Tool callingNo
Structured outputNot available
SpeedNot available
First tokenNot available
Maximum output15K
Catalog listing date2026-08-04