Text-Only Models with mm-ctx Vision Toolkit vs. Native Vision Models

Community Article
Published July 30, 2026

Summary

This report compares two types of AI agent setups on a multimodal task suite. Setup type 1 uses a text-only model with mm-ctx, a multimodal toolkit. Setup type 2 uses a model with native vision.

This report replaces an earlier version which used one run for each setup. Its aggregate claims did not agree with each other. Run-to-run variance changed its results. This version runs the full suite two times for each setup and we compute all numbers from the raw per-run results.

The benchmark has 23 tasks in three categories: retrieval, artifact creation, and file organization. Each task contains images, video, audio, PDFs, or a mix of these formats. Each setup does each task two times (23x2 case-runs). A deterministic checker and an LLM judge score each result. The correctness score is a 50/50 blend of the two scores. A timeout on either the setup or judge execution sets the score to zero. Thus timeouts decrease the correctness score.

We test four setups:

Setup Agent model Vision capable Toolkit backend
A gemini-3.5-flash Yes None
B qwen3.6-35b-a3b Yes None
C glm-5.2 + mm-ctx No gemini-3.5-flash-lite
D deepseek-v4-pro + mm-ctx No gemini-3.5-flash-lite

We access all models through OpenRouter and use the same agent harness (pi) for all four setups in non-interactive mode. The harness provides the agent with shell, file, and editing tools. In the toolkit setups, the agent uses mm-ctx CLI commands (find, peek, cat, grep, sql, wc) to inspect and extract content from binary and media files, while the toolkit backend model provides perception and vision capabilities. In the no-toolkit setups, the agent relies solely on its native capabilities. All results are evaluated by the same judge model (gemini-3.1-flash-lite), and each task is allotted a time budget of 360-600 seconds depending on its perceived difficulty.

Results

Correctness and speed numbers are means over two runs (23x2 case-runs per setup, @k=2). Completed is a count of finished case-runs, and a finished case-run is one with non-empty final answer and no timeout.

Overall scores

Rank Setup Mean correctness Mean speed Completed Timeouts Agent tokens Toolkit tokens
1 glm-5.2 + mm-ctx 88.4 100s 46/46 0 1.12M 3.13M
2 gemini-3.5-flash 83.0 150s 42/46 4 1.28M n/a
3 deepseek-v4-pro + mm-ctx 79.4 150s 44/46 2 1.12M 1.68M
4 qwen3.6-35b-a3b 44.3 119s 34/46 5 0.84M n/a

Per-run means (run 1 / run 2): glm-5.2 + mm-ctx 88.5 / 88.4, gemini-3.5-flash 78.6 / 87.5, deepseek-v4-pro + mm-ctx 72.1 / 86.6, qwen3.6-35b-a3b 38.6 / 50.0. Only the glm setup is stable across the two runs. The other three setups had a variance of 8.9-14.5 points.

Per-task results (mean of 2 runs)

Task Category Modality gemini-3.5 qwen3.6 glm + mm-ctx ds + mm-ctx
audio-claims-json artifact audio 71.3 75.0 42.5 61.3
audio-earnings-figure retrieval audio 100.0 100.0 100.0 50.0
bank-statement-reconcile artifact pdf 100.0 50.0 100.0 100.0
control-code-bug retrieval doc 100.0 60.0 77.5 77.5
control-config-grep retrieval doc 100.0 0.0 100.0 100.0
control-tiny-obvious retrieval image 100.0 0.0 100.0 100.0
cross-modal-order-match retrieval pdf, image 90.0 50.0 100.0 100.0
dedup-photos organization image 100.0 29.2 100.0 33.3
deep-pdf-crossref retrieval pdf 100.0 50.0 100.0 100.0
downloads-inventory-csv artifact image, video, audio, document 100.0 62.5 100.0 100.0
flag-medical-forms organization pdf 60.0 65.0 100.0 75.0
floorplan-room-report artifact image 0.0 40.0 50.0 50.0
invoices-to-csv artifact pdf, image 52.5 0.0 53.8 47.5
legal-clause-find retrieval pdf 100.0 90.0 90.0 90.0
ocr-needle-at-scale retrieval image 50.0 0.0 100.0 100.0
organize-by-subject organization image 100.0 100.0 100.0 100.0
paper-results-json artifact pdf 100.0 100.0 100.0 100.0
quarantine-non-floor-plans organization image, pdf 90.0 21.4 80.0 90.0
retrieve-sec-filing retrieval pdf 100.0 100.0 100.0 100.0
retrieve-video-product retrieval video 100.0 0.0 100.0 100.0
scanned-forms-extract-json artifact pdf 56.3 6.3 40.0 45.0
semantic-scene-retrieval retrieval image 90.0 0.0 100.0 50.0
video-timeline-json artifact video 50.0 18.8 100.0 56.3

Scores by task category

Category glm + mm-ctx gemini-3.5 ds + mm-ctx qwen3.6
Retrieval (11 tasks) 97.0 93.6 88.0 40.9
Organization (4 tasks) 95.0 87.5 74.6 53.9
Artifact (8 tasks) 73.3 66.3 70.0 44.1

Scores by data modality

Modality n glm + mm-ctx gemini-3.5 ds + mm-ctx qwen3.6
Audio 2 71.3 85.6 55.6 87.5
Document 2 88.8 100.0 88.8 30.0
Image 6 91.7 73.3 72.2 28.2
PDF 7 90.0 88.0 87.1 65.9
Video 2 100.0 75.0 78.1 9.4
Mixed 4 83.4 83.1 84.4 33.5

Single-modality groups contain only tasks with exactly one modality, while Mixed comprise tasks spanning multiple modalities: cross-modal-order-match, invoices-to-csv quarantine-non-floor-plans, and downloads-inventory-csv.

Key findings

1. glm-5.2 + mm-ctx (88.4) is the top setup. The only setup with zero timeouts. It has the lowest mean completion time (100s; gemini-3.5-flash has 150s). It is the only setup with a stable score across runs (88.5 / 88.4), having the top score on image-only (91.7), video-only (100.0), and PDF-only (90.0) tasks. It also has the top score in all three task categories.

2. gemini-3.5-flash surfered from timeouts. It timed out on four runs: floorplan-room-report (x2), ocr-needle-at-scale, and video-timeline-json. Each timeout receives a score of zero. Across the remaining 42 completed runs, its mean score is 91.0, slight higher than glm-5.2 + mm-ctx. While native perception of raw media is generally accurate, it is slower on bulk-image and long-video tasks. In contrast, the toolkit setup extracts only the relevant content, allowing the agent to complete the task within the allotted time. Since meeting the time limit is part of the benchmark, the penalty is appropriate. That said, the two setups are nearly identical in terms of pure correctness

3. deepseek-v4-pro + mm-ctx (79.4) exhibits high variance. Across the two runs, it scores 72.1 and 86.6, respectively. Its weakest performance comes on image organization (dedup-photos 33.3) and audio-only tasks (55.6). Despite using the same mm-ctx toolkit as glm-5.2, it scores nearly 9 points lower overall, indicating that the underlying model in the setup is as important as the toolkit.

4. qwen3.6-35b-a3b (44.3) produced multiple empty outputs. It failed to return a final answer on seven case-runs, five of them terminating in under 30 seconds. These include both runs of the two control tasks (control-config-grep, control-tiny-obvious). Since all other setups achieved a perfect score (100) on these controls, the failure appears to be model-specific rather than the benchmark. We plan to replace it with Inkling from Thinking Machines in future runs.

5. The toolkit increases token usage. glm-5.2 + mm-ctx uses 4.25M total tokens (1.12M agent + 3.13M toolkit backend), or about 3.3x the 1.28M used by gemini-3.5-flash. deepseek-v4-pro + mm-ctx uses 2.80M, or about 2.2x as many. The additional cost comes from the toolkit backend expectedly making vision requests for every binary file the agent processes. While mm-ctx employs custom pipelines and encoders for efficient token usage, the backend model (in this case, gemini-3.5-flash-lite) still contributes to the overall token consumption.

6. Text-only + mm-ctx setups are weakest on audio tasks. On the two audio-only tasks, gemini-3.5-flash (85.6) and qwen3.6 (87.5) outperform both toolkit setups (71.3 and 55.6). In particular, audio-claims-json is the only task where glm-5.2 + mm-ctx ranks last among all four setups, scoring (42.5).

Caveats

  • Two runs are a small sample. Three of the four setups vary by 8.9–14.5 points between runs. With only two runs per setup, per-task means remain noisy, and differences smaller than roughly 10 points should not be overinterpreted.
  • Timeouts impact scoring. A timeout in either the setup or the judge results in a score of zero for that run, which can substantially reduce the overall score.
  • Text-only + mm-ctx performance depends on the backend model. The toolkit's performance is tied to the capabilities of its backend multimodal model. A stronger backend model can improve scores on vision and audio tasks, while a weaker one can reduce them.

Verdict

On this evaluation, the text-only setup glm-5.2 + mm-ctx (88.4) outperforms the setup with native vision model gemini-3.5-flash (83.0). It completes every task within the allotted time and is, on average, 1.5x faster. On tasks completed by both setups, however, their scores are nearly identical (88.4 vs 91.0). The advantage of mm-ctx lies in efficient content extraction: by extracting only the relevant text from long media inputs, it keeps image-, video-, and PDF-heavy tasks within the allotted time. This comes at a cost, however: the toolkit consumes 3x more tokens and underperforms native vision on audio tasks.

Overall, these results suggest that a text-only model paired with a vision toolkit is a competitive approach for multimodal workloads. For directory-scale image, video, and PDF processing, it is the stronger option in this evaluation.


Next steps. Increase the number of runs per setup to at least five to reduce variance, replace qwen3.6-35b-a3b + pi setup with Inkling from Thinking Machines, and extend the benchmark with more difficult tasks in the dataset.

mm-ctx is an open-source project by VLM Run.

Community

Sign up or log in to comment