Matrena
Matrena is Russia’s strongest AI model. This page shows its results alongside Gemini 3.8 Flash, Opus 5, GigaChat 3.5, Alice, and GPT-5.6.
Explore resultsMatrena in numbers
Three results from the public comparison. The full table, charts, and evaluation notes follow below.
- SWE-bench Verified
- 79.0%
- Terminal-bench 2.1
- 91.4%
- τ2 Telecom
- 95.0%
What you can ask Matrena to do
Share the material or describe the task. Matrena works through it and helps you get a finished result.
Listens and watches
Works through audio, transcribes calls, and helps with video.
Finds information
Searches the web and gathers what the task needs.
Creates documents
Works with spreadsheets and Word, and assembles finished PDFs.
Writes scripts
Handles specific subtasks and helps automate work.
Beside other models
The highest published score in each row is marked. A hyphen means no data was published; it does not mean zero.
| Matrena | Gemini 3.8 Flash | Opus 5 | GigaChat 3.5 | Alice | GPT-5.6 | |
|---|---|---|---|---|---|---|
| SWE-bench VerifiedFixing issues in real projects | 79.0% | 80.0% | 96.0% | 42.6% | - | 82.2% |
| Terminal-bench 2.0Terminal tasks · version 2.0 | 89.0% | 89.4% | 89.1% | 13.5% | - | 91.9% |
| Terminal-bench 2.1Terminal tasks · version 2.1 | 91.4% | 89.4% | 89.1% | - | - | 88.8% |
| τ2 TelecomDialogue and tool use | 95.0% | - | - | 68.7% | - | - |
| MCP AtlasTasks with connected tools | 72.0% | - | - | - | - | - |
| OSWorldWork in graphical applications | 61.0% | 59.0% | 75.4% | - | - | 62.6% |
| GPQA DiamondDifficult natural-science questions | 88.1% | 95.3% | 93.2% | 61.1% | - | 94.6% |
Results on the same scale
Each chart shows one evaluation from 0 to 100%. Matrena’s place counts only models with published scores.
SWE-bench Verified
Matrena rank: 4 of 5Terminal-bench 2.0
Matrena rank: 4 of 5Terminal-bench 2.1
Matrena rank: 1 of 4τ2 Telecom
Matrena rank: 1 of 2MCP Atlas
Only Matrena has a published scoreOSWorld
Matrena rank: 3 of 4GPQA Diamond
Matrena rank: 4 of 5A hyphen in a chart means no published score is available.
Method
Matrena results come from our runs; other model results come from public reports. Run conditions may differ. The charts use the same numbers as the table. A hyphen means no publication, not 0%.
SWE-bench Verified
Real software project tasks: the agent reads an issue description and changes the code to fix it.
Terminal-bench
Terminal tasks include setup, data analysis, and file work. Versions 2.0 and 2.1 appear separately. The external 2.1 scores use the Terminus 2 agent comparison; GPT-5.6 here means Sol.
Using tools
τ2 Telecom tests dialogue and tool use in telecom tasks. MCP Atlas tests work with external tools through MCP.
OSWorld
Tasks in graphical computer applications: finding information, making changes, and producing a result.
GPQA Diamond
Difficult physics, chemistry, and biology questions written by subject experts.
ARC-AGI-2, MMMU, and MMMLU have no published scores for any model in this table.
Matrena