arxiv:2511.15552

Multimodal Evaluation of Russian-language Architectures

Published on Nov 19

· Submitted by

Ivan Sviridov on Nov 27

#1 Paper of the day

Upvote

Authors:

Ulyana Isaeva ,

Alexander Kharitonov ,

Vasily Konovalov ,

Elisei Rykov ,

Ivan Sviridov ,

Abstract

Mera Multi is an open multimodal evaluation framework for Russian-spoken architectures, addressing the lack of such benchmarks with 18 newly constructed tasks and a methodology to prevent benchmark leakage.

AI-generated summary

Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, and risks remain insufficiently understood. To address these issues, particularly in the context of the Russian language, where no multimodal benchmarks currently exist, we introduce Mera Multi, an open multimodal evaluation framework for Russian-spoken architectures. The benchmark is instruction-based and encompasses default text, image, audio, and video modalities, comprising 18 newly constructed evaluation tasks for both general-purpose models and modality-specific architectures (image-to-text, video-to-text, and audio-to-text). Our contributions include: (i) a universal taxonomy of multimodal abilities; (ii) 18 datasets created entirely from scratch with attention to Russian cultural and linguistic specificity, unified prompts, and metrics; (iii) baseline results for both closed-source and open-source models; (iv) a methodology for preventing benchmark leakage, including watermarking and licenses for private sets. While our current focus is on Russian, the proposed benchmark provides a replicable methodology for constructing multimodal benchmarks in typologically diverse languages, particularly within the Slavic language family.

View arXiv page View PDF Project page GitHub 13 Add to collection

Community

univanxx

Paper author Paper submitter 1 day ago

This work introduces MERA Multi, the first large-scale multimodal benchmark for Russian, encompassing 18 newly developed tasks across text, image, audio, and video, with a unified skill taxonomy, robust leakage protection (utilizing watermarking and membership inference), and a public leaderboard and codebase for evaluating both open and closed MLLMs.

blinoff

1 day ago

Good paper!

What is the most effective way to ensure that MLLMs for Russian achieve robust performance across image, audio, and video modalities, given the current performance gaps and cultural-linguistic specificity required?