Matreshka

Matrena

Matrena is Russia’s strongest AI model. This page shows its results alongside Gemini 3.8 Flash, Opus 5, GigaChat 3.5, Alice, and GPT-5.6.

Explore results

Matrena in numbers

Three results from the public comparison. The full table, charts, and evaluation notes follow below.

SWE-bench Verified
79.0%
Terminal-bench 2.1
91.4%
τ2 Telecom
95.0%

What you can ask Matrena to do

Share the material or describe the task. Matrena works through it and helps you get a finished result.

Listens and watches

Works through audio, transcribes calls, and helps with video.

Finds information

Searches the web and gathers what the task needs.

Creates documents

Works with spreadsheets and Word, and assembles finished PDFs.

Writes scripts

Handles specific subtasks and helps automate work.

Beside other models

The highest published score in each row is marked. A hyphen means no data was published; it does not mean zero.

Matrena compared with Gemini 3.8 Flash, Opus 5, GigaChat 3.5, Alice, and GPT-5.6
MatrenaGemini 3.8 FlashOpus 5GigaChat 3.5AliceGPT-5.6
SWE-bench VerifiedFixing issues in real projects79.0%80.0%96.0%42.6%-82.2%
Terminal-bench 2.0Terminal tasks · version 2.089.0%89.4%89.1%13.5%-91.9%
Terminal-bench 2.1Terminal tasks · version 2.191.4%89.4%89.1%--88.8%
τ2 TelecomDialogue and tool use95.0%--68.7%--
MCP AtlasTasks with connected tools72.0%-----
OSWorldWork in graphical applications61.0%59.0%75.4%--62.6%
GPQA DiamondDifficult natural-science questions88.1%95.3%93.2%61.1%-94.6%

Results on the same scale

Each chart shows one evaluation from 0 to 100%. Matrena’s place counts only models with published scores.

SWE-bench Verified

Matrena rank: 4 of 5
79.0%
Matrena
80.0%
Gemini 3.8 Flash
96.0%
Opus 5
42.6%
GigaChat 3.5
-
Alice
82.2%
GPT-5.6

Terminal-bench 2.0

Matrena rank: 4 of 5
89.0%
Matrena
89.4%
Gemini 3.8 Flash
89.1%
Opus 5
13.5%
GigaChat 3.5
-
Alice
91.9%
GPT-5.6

Terminal-bench 2.1

Matrena rank: 1 of 4
91.4%
Matrena
89.4%
Gemini 3.8 Flash
89.1%
Opus 5
-
GigaChat 3.5
-
Alice
88.8%
GPT-5.6

τ2 Telecom

Matrena rank: 1 of 2
95.0%
Matrena
-
Gemini 3.8 Flash
-
Opus 5
68.7%
GigaChat 3.5
-
Alice
-
GPT-5.6

MCP Atlas

Only Matrena has a published score
72.0%
Matrena
-
Gemini 3.8 Flash
-
Opus 5
-
GigaChat 3.5
-
Alice
-
GPT-5.6

OSWorld

Matrena rank: 3 of 4
61.0%
Matrena
59.0%
Gemini 3.8 Flash
75.4%
Opus 5
-
GigaChat 3.5
-
Alice
62.6%
GPT-5.6

GPQA Diamond

Matrena rank: 4 of 5
88.1%
Matrena
95.3%
Gemini 3.8 Flash
93.2%
Opus 5
61.1%
GigaChat 3.5
-
Alice
94.6%
GPT-5.6

A hyphen in a chart means no published score is available.

Method

Matrena results come from our runs; other model results come from public reports. Run conditions may differ. The charts use the same numbers as the table. A hyphen means no publication, not 0%.

SWE-bench Verified

Real software project tasks: the agent reads an issue description and changes the code to fix it.

Terminal-bench

Terminal tasks include setup, data analysis, and file work. Versions 2.0 and 2.1 appear separately. The external 2.1 scores use the Terminus 2 agent comparison; GPT-5.6 here means Sol.

Using tools

τ2 Telecom tests dialogue and tool use in telecom tasks. MCP Atlas tests work with external tools through MCP.

OSWorld

Tasks in graphical computer applications: finding information, making changes, and producing a result.

GPQA Diamond

Difficult physics, chemistry, and biology questions written by subject experts.

ARC-AGI-2, MMMU, and MMMLU have no published scores for any model in this table.

Matrena

Matrena is Russia’s strongest AI model.