AI & ML interests

None defined yet.

Recent Activity

Articles

nwaughachukwuma 
posted an update 5 days ago
view post
Post
103
Here's how the best OCR/chat models on vlmrun gateway - https://vlm.run/gateway - performed on a very hard (x2) to read handwritten French weather-log document.

Scores are 0–100, averaged over 3 runs and two LLM judges - Claude Opus 5.5 and GPT-5.6 Sol. Latency is average time per request.
  • 1 reply
·
nwaughachukwuma 
posted an update 16 days ago
view post
Post
69
Cheap, fast, and efficient OCR + RAG pipeline with VLM Run Gateway and TypeSafe AI Jev.
nwaughachukwuma 
posted an update about 1 month ago
view post
Post
3235
It’s easy to get distracted by benchmarks, throughput (tok/s), and all the hype around frontier model releases.

This is Shiny Model Syndrome, which makes engineers and teams forget the basic physics of production software, i.e., using the right tool for the job and optimizing for ease of integration.

- Teams spend huge amounts of money on frontier models for document parsing, OCR, detection, segmentation, and other task-specific visual AI workflows.

- Inference marketplaces don’t find it profitable to list task-specific models like glm-ocr, paddleocr, or dots.mocr, even though they’re all superior to frontier VLMs for document parsing and OCR.

- Engineers stitch together multiple endpoints for different use cases across the long tail of visual AI. Those who choose to self-host instead deal with painful infrastructure and GPU ops.

At VLM Run, we wanted one place to run OCR models, VLMs, and ViTs that we could confidently use for our own internal agents and evals. The gateway was born out of that need, and we’ve since opened it to the public.

The gateway exposes a single OpenAI-compatible endpoint for the long tail of visual AI across OCR, document parsing, VQA, detection, segmentation, embeddings, and transcription. Simply point the base_url of your OpenAI SDK at gateway.vlm.run/v1/openai, or ask your agent to connect via MCP (gateway.vlm.run/mcp).

You can swap the model name to compare glm-ocr, dots.mocr, paddleocr-vl-1.6, qwen3.8-27b, gemma4-26b-a4b, and more. We handle serving, runtime, and pipelining behind the scenes to give you high-quality visual intelligence.

- https://vlm.run/gateway
- https://huggingface.co/blog/vlm-run/introducing-gateway
- https://www.vlm.run/blog/introducing-gateway
spillai 
posted an update about 1 month ago
view post
Post
3211
We're excited to introduce VLM Run Gateway - a new unified OpenAI-compatible API for running open-weight VLMs, OCR VLMs and ViT-based vision models.

https://vlm.run/gateway
Full model catalog: https://vlm.run/gateway/models
Blog post announcement: https://www.vlm.run/blog/introducing-gateway

Try different models on the gateway simply by updating the model name. Free to use and no sign-up required for now (in alpha).

$ uvx vlmrun gw models
$ uvx vlmrun gw chat <doc>.pdf -m glm-ocr
$ uvx vlmrun gw chat <doc>.pdf -m deepseek-ocr-2
$ uvx vlmrun gw chat <doc>.pdf -m pp-ocrv6
$ uvx vlmrun gw chat <img>.jpg -m qwen/qwen3.5-0.8b -p "describe the image"
$ uvx vlmrun gw chat <vid>.mp4 -m qwen/qwen3.5-0.8b -p "describe the video"

spillai 
published an article about 1 month ago
view article
Article

VLM Run Gateway: Run open-weight OCR, VLM and vision models behind one API

vlm-run
•
• 6
nwaughachukwuma 
posted an update about 1 month ago
view post
Post
79
In 1887, Eadweard Muybridge lined up twelve cameras to answer one question: how does a body actually move?

Today, you can easily tell by pointing the video at @vlmrun Gateway.

$ uvx vlmrun gw chat clip.mp4 -m vitpose-plus-large --method pose

One call. 612 keypoints. 36 frames. $0.000998.
nwaughachukwuma 
posted an update about 2 months ago
view post
Post
95
You can now use paddleocr-vl on VLM Run Gateway.

Via the CLI:
uvx vlmrun gw models
uvx vlmrun config set --api-key '<VLMRUN_API_KEY>' # anon-user, rate-limited
uvx vlmrun gw chat <doc>.pdf -m paddleocr/pp-ocrv6

uvx vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr
uvx vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr --json-mode
uvx vlmrun gw chat <doc>.pdf -m deepseek-ai/deepseek-ocr-2
uvx vlmrun gw chat <doc>.pdf -m rednote-hilab/dots.mocr


OpenAI SDK:
from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.vlm.run/v1/openai",
    api_key="<VLMRUN_API_KEY>",
)

response = client.chat.completions.create(
    model="paddlepaddle/paddleocr-vl",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {
                        "url": "https://storage.googleapis.com/vlm-data-public-prod/hub/examples/document.receipt/playground/2.jpg"
                    },
                },
            ],
        }
    ],
    extra_body={"method": "table"},
)

print(response.choices[0].message.content)


or Curl:
curl https://gateway.vlm.run/v1/openai/chat/completions \
  -X POST \
  -H "Authorization: Bearer <VLMRUN_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "paddlepaddle/paddleocr-vl",
    "method": "table",
    "messages": [
      {
        "role": "user",
        "content": [
          {
            "type": "image_url",
            "image_url": {
              "url": "https://storage.googleapis.com/vlm-data-public-prod/hub/examples/document.receipt/playground/2.jpg"
            }
          }
        ]
      }
    ]
  }'


Docs: https://docs.vlm.run/gateway

Catalog: https://docs.vlm.run/gateway/models

MCP: https://docs.vlm.run/gateway/mcp-server

Colab Quickstart: https://colab.research.google.com/drive/1RkuVIyuc5Po-UlcSlFyJCam5tjCm9IHM?usp=sharing
nwaughachukwuma 
posted an update about 2 months ago
view post
Post
3587
# VLM Run Gateway: Run GLM-OCR, DeepSeek-OCR-2, and Dots.mocr with an OpenAI Compatible API

Open-weight OCR VLMs have advanced significantly over the past year, yet most teams still rely on frontier VLMs for document parsing because researching, evaluating, and deploying the right models remains challenging.

So we built VLM Run Gateway: one OpenAI-compatible endpoint for open-weight OCR and VLM models.

If you’re using frontier VLMs primarily for OCR/document parsing, open-weight OCR models can be dramatically cheaper and often very accurate. With a one-line change, you can switch between open-weight OCR VLMs (DeepSeek OCR 2, GLM-OCR, dots.mocr, Paddle OCR VL, PP-OCRv6, etc.) and process 100K+ pages for under $60.

Try it out quickly via the CLI:
uvx vlmrun gw models
uvx vlmrun config set --api-key '<VLMRUN_API_KEY>' # anon-user, rate-limited
uvx vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr

uvx vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr --json-mode
uvx vlmrun gw chat <doc>.pdf -m deepseek-ai/deepseek-ocr-2
uvx vlmrun gw chat <doc>.pdf -m rednote-hilab/dots.mocr
uvx vlmrun gw chat <doc>.pdf -m paddleocr/pp-ocrv6


OpenAI SDK:
client = OpenAI(
    base_url="https://gateway.vlm.run/v1/openai",
    api_key="<VLMRUN_API_KEY>",
)

response = client.chat.completions.create(
    model="rednote-hilab/dots.mocr",
    messages=[{
      "role": "user",
		  "content": [{
            "type": "document_url",
            "document_url": {"url": "https://.../invoice.pdf"},
      }],
    }],
    extra_body={"document_dpi": 72},
)


Docs: https://docs.vlm.run/gateway

Catalog: https://docs.vlm.run/gateway/models

MCP: https://docs.vlm.run/gateway/mcp-server

Colab Quickstart: https://colab.research.google.com/drive/1RkuVIyuc5Po-UlcSlFyJCam5tjCm9IHM?usp=sharing

Read the full post here: https://huggingface.co/blog/vlm-run/intro-to-vlmrun-gateway
nwaughachukwuma 
published an article about 2 months ago
view article
Article

VLM Run Gateway: Run GLM-OCR, DeepSeek-OCR-2, dots.mocr with an OpenAI Compatible API

vlm-run
•
• 6
nwaughachukwuma 
posted an update about 2 months ago
view post
Post
2727
Following this post - https://huggingface.co/posts/nwaughachukwuma/366831808712777, I made another run on a fresh RTX PRO 4000 using inkling, gemini-3.5-flash, and glm-5.2+mm-ctx

| case | gemini-3.5-flash @k=1 | glm-5.2+mm @k=1 | inkling @k=1 |
| --- | --- | --- | --- |
| **MEAN correctness** | **78.2** | **79.5** | **57.9** |
| **MEAN speed (s)** | **111** | **123** | **69** |
| **completed (case-runs)** | **20/23** | **21/23** | **18/23** |
nwaughachukwuma 
published an article 2 months ago
view article
Article

Text-Only Models with mm-ctx Vision Toolkit vs. Native Vision Models

vlm-run
•
nwaughachukwuma 
posted an update 2 months ago
view post
Post
687
Can a text-only model + a vision toolkit (mm-ctx) match a native vision model?

We benchmarked 4 setups on 23 multimodal tasks (image, video, audio, PDF):

• glm-5.2 (text-only) + mm-ctx: 88.4
• gemini-3.5-flash (vision): 83
• deepseek-v4-pro (text-only) + mm-ctx: 79.4
• qwen3.6-35b-a3b (vision): 44.3

The best text-only setup glm-5.2 + mm outperformed gemini-3.5-flash, the top vision model, by 5.4 points (6.5%). It was also:

• 1.5x faster (100s vs 150s mean per task)
• the only setup with zero timeouts (46/46 completed; gemini timed out 4x on bulk-image and long-video tasks)
• the only setup stable across runs (88.5 / 88.4)
• top on video (100.0), image (91.7), and PDF (90.0) tasks

The trade-offs: the toolkit consumed 3.3x more tokens (4.25M vs 1.28M), and lost on audio (85.6 vs 71.3).

On completed tasks alone the two are nearly identical (91.0 vs 88.4): the toolkit's edge is efficient extraction that keeps long media tasks inside the time budget.

Full report: https://huggingface.co/blog/vlm-run/text-only-models-with-mm
  • 8 replies
·
nwaughachukwuma 
posted an update 3 months ago
view post
Post
748
# One API for Every Visual & OCR Models.

The VLM Run Gateway is a fully compatible API for OpenAI chat completions for visual intelligence. If you’re building document extraction or visual understanding, the Gateway exposes OCR, VQA, and detection behind a single interface you already know.

Read the docs: https://docs.vlm.run/gateway/introduction.

We actively support the following recent OCR and VQA models, which you can try today at no cost:

* zai-org/glm-ocr
* rednote-hilab/dots.mocr
* paddleocr/pp-ocrv6
* qwen/qwen3.5-0.8b

## Quickstart

### CLI
uvx vlmrun gw models
uvx vlmrun config set --api-key '<VLMRUN_API_KEY>' # anon-user, rate-limited
uvx vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr

uvx vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr --json-mode
uvx vlmrun gw chat <doc>.pdf -m deepseek-ai/deepseek-ocr-2
uvx vlmrun gw chat <doc>.pdf -m rednote-hilab/dots.mocr
uvx vlmrun gw chat <doc>.pdf -m paddleocr/pp-ocrv6


### OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.vlm.run/v1/openai",
    api_key="<VLMRUN_API_KEY>",
)


response = client.chat.completions.create(
    model="zai-org/glm-ocr",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "document_url",
                    "document_url": {
                        "url": "https://storage.googleapis.com/vlm-data-public-prod/hub/examples/finance.sec-filings/tsla-8k.pdf"
                    },
                },
            ],
        }
    ],
    extra_body={"method": "markdown", "document_dpi": 150},
)

print(response.choices[0].message.content)


## Auth and limits

Anonymous auth is enabled, so you can omit the authorization header entirely, or send Bearer "" or Bearer vlmrun. Rate limits are 60 req/min and 1000 req/hr.
spillai 
posted an update 5 months ago
view post
Post
8800
mm-ctx – fast, multimodal context for agents.

LLM-based agents handle text incredibly well, but images, videos, or PDFs with visual content are hard to interpret. mm-ctx gives your CLI agent multi-modal skills.

Try it interactively in Spaces: vlm-run/mm-ctx

Readme: https://vlm-run.github.io/mm/
PyPI: https://pypi.org/project/mm-ctx
SKILL.md: https://github.com/vlm-run/skills/blob/main/skills/mm-cli-skill/SKILL.md

mm-ctx is meant to feel familiar: the UNIX tools we already love (find/cat/grep/wc), rebuilt for file types LLMs can't read natively and designed to work with agents via the CLI.
- mm grep "invoice #1234" ~/Downloads searches across PDFs and returns line-numbered matches
- mm cat <document>.pdf returns a metadata description of the file
- mm cat <photo>.jpg returns a caption of the photo
- mm cat <video>.mp4 returns a caption of the video

A few things we obsessed over:
⚡ Speed: Rust core for the hot paths
🏠 Local-first, BYO model: Uses any OpenAI-compatible endpoint: Ollama, vLLM/SGLang, LMStudio with any multimodal LLM (Gemma4, Qwen3.5, GLM-4.6V).
🔗 Composable: stdin + structured outputs
🤖 Drops into any agent via mm-cli-skills: Claude Code, Codex, Gemini CLI, OpenClaw.

We’d love to hear your feedback! Especially on the CLI and what file types and workflows you would like to see next.
  • 2 replies
·