AV LLMs
A collection of Audio, Video and Visual LLMs.
-
Text-to-Speech • Updated • 494 -
OpenVoice
🤗1.14kGenerate speech in a cloned voice from a short audio clip
-
dataautogpt3/ProteusV0.3
Text-to-Image • Updated • 18.6k • 96 -
ByteDance/SDXL-Lightning
Text-to-Image • Updated • 105k • 2.2k -
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 3.99M • • 6.57k -
stabilityai/TripoSR
Image-to-3D • Updated • 234k • 672 -
Efficient-Large-Model/VILA-7b
Text Generation • 7B • Updated • 99 • 27 -
google/paligemma-3b-pt-896
Image-Text-to-Text • 3B • Updated • 331 • 125 -
microsoft/Phi-3-vision-128k-instruct
Text Generation • 4B • Updated • 51.5k • 973 -
stabilityai/stable-audio-open-1.0
Text-to-Audio • 1B • Updated • 19.1k • 1.7k -
OpenVLA: An Open-Source Vision-Language-Action Model
Paper • 2406.09246 • Published • 47 -
aiola/whisper-medusa-v1
2B • Updated • 88 • 180 -
merve/idefics3llama-vqav2
Updated • 8 -
black-forest-labs/FLUX.1-schnell
Text-to-Image • 12B • Updated • 649k • • 6.16k -
Llama3.1 S V0.2 Checkpoint 2024 08 20
😻115Chat with Llama3.1 using spoken audio or synthesize speech
-
gpt-omni/mini-omni
Text-to-Speech • Updated • 445 -
fishaudio/fish-speech-1.4
Text-to-Speech • Updated • 248 • 460 -
Tonic's GOT OCR
📲178GOT - OCR (from : UCAS, Beijing)
-
stepfun-ai/GOT-OCR2_0
Image-Text-to-Text • 0.7B • Updated • 562k • 1.56k -
apple/coreml-sam2-large
Mask Generation • Updated • 48 • 35 -
coreml-projects/sam-2-studio
Updated • 28 -
mistralai/Pixtral-12B-2409
Updated • 8.71k • 700 -
allenai/Molmo-72B-0924
Image-Text-to-Text • 73B • Updated • 10.8k • 299 -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.35M • • 3.42k -
Revai/reverb-asr
Automatic Speech Recognition • Updated • 10 • 93 -
GOT Online
💬361Extract and format text from images with advanced OCR
-
facebook/vfusion3d
Image-to-3D • 0.5B • Updated • 45 • 65 -
facebook/cotracker
Updated • 1.17k • 38 -
rhymes-ai/Aria
Image-Text-to-Text • 25B • Updated • 41.7k • 639 -
SWivid/F5-TTS
Text-to-Speech • Updated • 816k • 1.21k -
Ichigo Llama3.1 S Instruct
🏢64Generate text from audio recordings
-
kyutai/moshiko-mlx-q4
Updated • 1.7k • 30 -
kyutai/moshiko-mlx-q8
Updated • 474 • 7 -
Open VLM Video Leaderboard
🌎136VLMEvalKit Eval Results in video understanding benchmark
-
jimmycarter/LibreFLUX
Text-to-Image • 12B • Updated • 198 • 172 -
microsoft/OmniParser
Image-Text-to-Text • Updated • 701 • 1.72k -
Aya Models
🌍336Interact with the Aya family of models.
-
CohereLabs/aya-expanse-32b
Text Generation • 32B • Updated • 7.12k • • 303 -
stabilityai/stable-diffusion-3.5-medium
Text-to-Image • 2B • Updated • 90.5k • • 1.18k -
OuteAI/OuteTTS-0.1-350M
Text-to-Speech • 0.4B • Updated • 51 • 304 -
vidore/colpali
Visual Document Retrieval • Updated • 2.01k • 488 -
vidore/colpali-v1.2
Visual Document Retrieval • Updated • 46.5k • 113 -
si-pbc/hertz-dev
Audio-to-Audio • Updated • 216 -
Talk To Ultravox
⚡38Talk to Fixie.ai's Ultravox with WebRTC ⚡️
-
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Paper • 2411.10440 • Published • 133 -
Xkev/Llama-3.2V-11B-cot
Image-Text-to-Text • 11B • Updated • 825 • 158 -
google/paligemma-3b-pt-224
Image-Text-to-Text • 3B • Updated • 240k • 619 -
apple/coreml-mobileclip
Updated • 2.57k • 55 -
InstantX/InstantIR
Image-to-Image • Updated • 1 • 179 -
InstantIR
🖼85diffusion-based Image Restoration model
-
Flux IP Adapter
🖼170Prompt with Images in flux[dev]
-
Image Preferences - Argilla annotation space
🖼39A community project to create an image preferences dataset.
-
fishaudio/fish-speech-1.5
Text-to-Speech • Updated • 3.48k • 777 -
meta-llama/Llama-3.3-70B-Instruct
Text Generation • 71B • Updated • 366k • • 3.11k -
Paligemma2 Vqav2
🐨48PaliGemma2 LoRA finetuned on VQAv2
-
VisionZip: Longer is Better but Not Necessary in Vision Language Models
Paper • 2412.04467 • Published • 119 -
fancyfeast/llama-joycaption-alpha-two-hf-llava
8B • Updated • 23.9k • 218 -
taohu/mask
Updated • 5 -
[MASK] is All You Need
Paper • 2412.06787 • Published • 2 -
Open VLM Leaderboard
🌎1.03kVLMEvalKit Evaluation Results Collection
-
microsoft/LLM2CLIP-Llama3.2-1B-EVA02-L-14-336
Zero-Shot Image Classification • Updated • 12 -
LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation
Paper • 2411.04997 • Published • 39 -
Generative Powers of Ten
Paper • 2312.02149 • Published • 7 -
StoryStar
💬24Fantasy story generator
-
GoodiesHere/Apollo-LMMs-Apollo-7B-t32
Video-Text-to-Text • Updated • 49 • 58 -
Apollo: An Exploration of Video Understanding in Large Multimodal Models
Paper • 2412.10360 • Published • 148 -
Qwen/Qwen2-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 630k • 1.29k -
XiaoduoAILab/Xmodel_VLM
Text Generation • 2B • Updated • 284 • 13 -
nvidia/Cosmos-1.0-Diffusion-14B-Text2World
Updated • 16 • 61 -
nvidia/Cosmos-1.0-Autoregressive-12B
Updated • 25 • 31 -
nvidia/Cosmos-1.0-Autoregressive-13B-Video2World
Updated • 7 • 33 -
nvidia/Cosmos-1.0-Diffusion-7B-Text2World
Text-to-Video • Updated • 1.84k • 236 -
nvidia/Cosmos-1.0-Diffusion-14B-Video2World
Updated • 1.25k • 61 -
Stable Point-Aware 3D
⚡470Generate 3D models from images
-
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.1M • • 7.15k -
Kokoro TTS
❤3.5kUpgraded to v1.0!
-
openbmb/MiniCPM-o-2_6
Any-to-Any • 9B • Updated • 212k • 1.3k -
TTS Spaces Arena
🤗488Blind vote on HF TTS models!
-
google/paligemma2-10b-pt-896
Image-Text-to-Text • 10B • Updated • 92 • 34 -
NovaSky-AI/Sky-T1-32B-Preview
Text Generation • 33B • Updated • 215 • • 549 -
MiniMaxAI/MiniMax-VL-01
Image-Text-to-Text • 456B • Updated • 30.3k • 286 -
SmolVLM
📊68Generate descriptions from images and text prompts
-
HKUSTAudio/Llasa-3B
Text-to-Speech • 4B • Updated • 318 • 531 -
HuggingFaceTB/SmolVLM-500M-Instruct
Image-Text-to-Text • 0.5B • Updated • 136k • 199 -
deepseek-ai/Janus-Pro-7B
Any-to-Any • Updated • 10.9k • 3.7k -
Kokoro TTS Zero
🎴312✨[With v1.0.0] Accelerated TTS on Kokoro-82M
-
kyutai/hibiki-2b-mlx-bf16
Translation • Updated • 33 • 23 -
kyutai/hibiki-2b-pytorch-bf16
Translation • Updated • 236 • 63 -
ARTPARK-IISc/Vaani
Viewer • Updated • 9.19M • 22.4k • 160 -
Zyphra/Zonos-v0.1-hybrid
Text-to-Speech • 2B • Updated • 1.12k • • 1.11k -
Zyphra/Zonos-v0.1-transformer
Text-to-Speech • 2B • Updated • 41.3k • • 436 -
microsoft/OmniParser-v2.0
Updated • 3.2k • 1.36k -
Paligemma2 Mix
🌖98Generate text answers or segment objects from images
-
google/paligemma2-3b-mix-448
Image-Text-to-Text • 3B • Updated • 10k • 69 -
google/paligemma2-3b-mix-224
Image-Text-to-Text • 3B • Updated • 22.4k • 65 -
google/paligemma2-28b-mix-224
Image-Text-to-Text • 28B • Updated • 11 • 5 -
google/paligemma2-28b-mix-448
Image-Text-to-Text • 28B • Updated • 95 • 29 -
google/paligemma2-10b-mix-224
Image-Text-to-Text • 10B • Updated • 157 • 10 -
google/paligemma2-10b-mix-448
Image-Text-to-Text • 10B • Updated • 2.13k • 37 -
stepfun-ai/stepvideo-t2v
Text-to-Video • 29B • Updated • 77 • 480 -
stepfun-ai/stepvideo-t2v-turbo
29B • Updated • 99 -
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Paper • 2502.10248 • Published • 57 -
HuggingFaceTB/SmolVLM2-2.2B-Instruct
Image-Text-to-Text • 2B • Updated • 143k • 335 -
nvidia/canary-1b
Automatic Speech Recognition • Updated • 3.21k • 459 -
Wan-AI/Wan2.1-I2V-14B-720P
Image-to-Video • 16B • Updated • 20.1k • • 611 -
fastrtc/kokoro-onnx
Updated • 13 -
Fastphone
🐠2Download and run a Hugging Face app
-
microsoft/Phi-4-multimodal-instruct
Automatic Speech Recognition • 6B • Updated • 198k • 1.62k -
microsoft/Magma-8B
Robotics • 9B • Updated • 1k • 418 -
Magma UI
📚47Magma-8B model for UI Agents
-
Di♪♪Rhythm
🎶689Blazingly Fast and Embarrassingly Simple Song Generation
-
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Paper • 2503.01183 • Published • 29 -
ASLP-lab/DiffRhythm-vae
Updated • 43 -
ASLP-lab/DiffRhythm-base
Updated • 125 • 171 -
Large Language Diffusion Models
Paper • 2502.09992 • Published • 129 -
GSAI-ML/LLaDA-8B-Instruct
Text Generation • 8B • Updated • 435k • 364 -
unsloth/gemma-3-12b-pt
Image-Text-to-Text • 12B • Updated • 10.4k • • 5 -
google/gemma-3-27b-it
Image-Text-to-Text • 27B • Updated • 417k • • 2.05k -
sesame/csm-1b
Text-to-Speech • 2B • Updated • 133k • • 2.47k -
unsloth/gemma-3-27b-it-GGUF
Image-Text-to-Text • 27B • Updated • 38.8k • 206 -
docling-project/SmolDocling-256M-preview
Image-Text-to-Text • 0.3B • Updated • 26.5k • 1.62k -
starvector/starvector-8b-im2svg
Text Generation • 8B • Updated • 1.94k • 571 -
starvector/starvector-1b-im2svg
Text Generation • 1B • Updated • 3.64k • 199 -
Tokenize Image as a Set
Paper • 2503.16425 • Published • 16 -
kyutai/moshika-vis-pytorch-bf16
9B • Updated • 59 -
kyutai/Babillage
Viewer • Updated • 465k • 539 • 14 -
ByteDance/InfiniteYou
Text-to-Image • Updated • 892 • 645 -
InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
Paper • 2503.16418 • Published • 36 -
openfree/flux-chatgpt-ghibli-lora
Text-to-Image • Updated • 479 • • 510 -
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
Paper • 2504.00595 • Published • 38 -
weizhiwang/Open-Qwen2VL
Image-Text-to-Text • Updated • 108 • 21 -
ostris/Flex.1-alpha-Redux
Text-to-Image • Updated • 227 • 116 -
unsloth/Llama-4-Scout-17B-16E-Instruct-unsloth-bnb-4bit
Image-Text-to-Text • 112B • Updated • 595 • 80 -
unsloth/Llama-4-Scout-17B-16E-Instruct-unsloth-bnb-8bit
Image-Text-to-Text • 109B • Updated • 32 • 9 -
SmolVLM: Redefining small and efficient multimodal models
Paper • 2504.05299 • Published • 212 -
canopylabs/3b-hi-ft-research_release
Text-to-Speech • 3B • Updated • 125 • 26 -
canopylabs/3b-es_it-ft-research_release
Text-to-Speech • 3B • Updated • 620 • 22 -
nvidia/C-RADIOv2-g
Image Feature Extraction • 1B • Updated • 149 • 13 -
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Paper • 2504.10479 • Published • 312 -
OpenGVLab/InternVL3-1B
Image-Text-to-Text • 0.9B • Updated • 104k • 85 -
OpenGVLab/InternVL3-78B
Image-Text-to-Text • 78B • Updated • 8.17k • 239 -
InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
Paper • 2504.05303 • Published • 5 -
Dia 1.6B
👯1.79kGenerate realistic dialogue from a script, using Dia!
-
nari-labs/Dia-1.6B
Text-to-Speech • 2B • Updated • 27.5k • • 2.92k -
Describe Anything: Detailed Localized Image and Video Captioning
Paper • 2504.16072 • Published • 66 -
nvidia/DAM-3B-Self-Contained
Image-Text-to-Text • Updated • 4.5k • 25 -
nvidia/DAM-3B-Video
Image-Text-to-Text • Updated • 3.15k • 59 -
nvidia/DAM-3B
Image-Text-to-Text • Updated • 16k • 130 -
Qwen/Qwen2.5-Omni-3B
Any-to-Any • 6B • Updated • 294k • 359 -
MMaDA: Multimodal Large Diffusion Language Models
Paper • 2505.15809 • Published • 99 -
One RL to See Them All: Visual Triple Unified Reinforcement Learning
Paper • 2505.18129 • Published • 63 -
PlayDiffusion
🎨119Generate modified audio from text and voice
-
lerobot/smolvla_base
Robotics • 0.5B • Updated • 109k • 459 -
stockmark/Stockmark-2-VL-100B-beta
Image-Text-to-Text • 96B • Updated • 171 • 23 -
Qwen/Qwen2.5-Omni-7B
Any-to-Any • 11B • Updated • 316k • 1.96k -
Qwen2.5-Omni Technical Report
Paper • 2503.20215 • Published • 174 -
Chatterbox TTS
🍿1.76kExpressive Zeroshot TTS
-
ResembleAI/chatterbox
Text-to-Speech • Updated • 1.63M • • 1.83k -
ByteDance/Dolphin
Image-Text-to-Text • 0.4B • Updated • 615 • 519 -
nanonets/Nanonets-OCR-s
Image-Text-to-Text • 4B • Updated • 58.3k • 1.6k -
Nanonets Ocr S
👁35https://nanonets.com/research/nanonets-ocr-s/
-
calcuis/cosmos-predict2-gguf
Text-to-Image • 14B • Updated • 7.4k • 33 -
Arrexel/pattern-diffusion
Text-to-Image • 0.9B • Updated • 89 • 115 -
numind/NuMarkdown-8B-Thinking
Image-to-Text • 8B • Updated • 183k • 496 -
Qwen/Qwen-Image
Text-to-Image • 20B • Updated • 258k • • 2.67k -
dots-studio/dots.ocr
Image-Text-to-Text • 3B • Updated • 1.13M • 1.33k -
Runware/Qwen-Image-Edit
Image-to-Image • 20B • Updated • 86 • 18 -
Qwen Image Edit
✒868Edit images using natural language instructions
-
Qwen/Qwen-Image-Edit
Image-to-Image • 20B • Updated • 110k • • 2.54k -
zju-community/matchanything_eloftr
16.1M • Updated • 7.13k • 84 -
MatchAnything
🏢262Find similar images and match them across collections
-
microsoft/VibeVoice-1.5B
Text-to-Speech • 3B • Updated • 712k • 2.49k -
bytedance-research/USO
Text-to-Image • Updated • 228 • 191 -
FastVLM WebGPU
🍎450Real-time video captioning powered by FastVLM
-
onnx-community/FastVLM-0.5B-ONNX
Image-Text-to-Text • Updated • 645 • 109 -
apple/FastVLM-0.5B
Text Generation • 0.8B • Updated • 5.77k • 406 -
Qwen/Qwen3-Omni-30B-A3B-Instruct
Any-to-Any • 35B • Updated • 662k • 1.01k -
smolagents/SmolVLM2-2.2B-Instruct-Agentic-GUI
Image-Text-to-Text • 2B • Updated • 333 • 68 -
facebook/sam2.1-hiera-large
Mask Generation • 0.2B • Updated • 259k • 147 -
PaddlePaddle/PaddleOCR-VL
Image-Text-to-Text • 1.0B • Updated • 7.86k • 1.67k -
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Paper • 2510.14528 • Published • 129 -
PaddleOCR-VL Online Demo
📈255Extract text, tables, formulas, and charts from images
-
nanonets/Nanonets-OCR2-3B
Image-Text-to-Text • 4B • Updated • 269k • 516 -
nanonets/Nanonets-OCR2-1.5B-exp
Image-Text-to-Text • 2B • Updated • 299 • 51 -
deepseek-ai/DeepSeek-OCR
Image-Text-to-Text • 3B • Updated • 2.08M • • 3.39k -
lightonai/LightOnOCR-1B-1025
Image-to-Text • 1B • Updated • 29.1k • 256 -
Qwen Image Edit Camera Control
🎬2.27kFast 4 step inference with Qwen Image Edit 2509
-
depth-anything/DA3NESTED-GIANT-LARGE
Depth Estimation • 2B • Updated • 29k • 54 -
microsoft/Fara-7B
Image-Text-to-Text • 8B • Updated • 8.73k • 621 -
tencent/HunyuanOCR
Image-Text-to-Text • 1B • Updated • 929k • 827 -
ApsaraStackMaaS/EvoQwen2.5-VL-Retriever-7B-v1
Visual Document Retrieval • 8B • Updated • 345 • 18 -
apple/starflow
Updated • 286 -
microsoft/VibeVoice-Realtime-0.5B
Text-to-Speech • 1B • Updated • 190k • 1.3k -
ServiceNow-AI/Apriel-1.5-15b-Thinker
Image-Text-to-Text • 15B • Updated • 642 • 475 -
zai-org/GLM-4.6V-Flash
Image-Text-to-Text • 10B • Updated • 110k • • 628 -
Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
Paper • 2512.15603 • Published • 71 -
Qwen/Qwen-Image-Edit-2511
Image-to-Image • 20B • Updated • 255k • • 1.45k -
microsoft/TRELLIS.2-4B
Image-to-3D • Updated • 1.88M • 1.29k -
inclusionAI/TwinFlow-Z-Image-Turbo
Text-to-Image • Updated • 49 • 213 -
Phr00t/Qwen-Image-Edit-Rapid-AIO
Text-to-Image • Updated • 2.55k -
LiquidAI/LFM2.5-Audio-1.5B
Audio-to-Audio • 1B • Updated • 5.05k • 469 -
LiquidAI/LFM2.5-VL-1.6B
Image-Text-to-Text • 2B • Updated • 30.6k • 330 -
black-forest-labs/FLUX.2-klein-9B
Image-to-Image • 9B • Updated • 200k • • 1.62k -
black-forest-labs/FLUX.2-dev
Image-to-Image • 32B • Updated • 405k • • 2.38k -
nvidia/personaplex-7b-v1
Audio-to-Audio • 8B • Updated • 183k • • 2.79k -
bevaya/GutenOCR-3B
Image-Text-to-Text • 4B • Updated • 353 • 27 -
bevaya/GutenOCR-7B
Image-Text-to-Text • 8B • Updated • 142 • 26 -
microsoft/VibeVoice-ASR
Automatic Speech Recognition • 9B • Updated • 723k • 1.3k -
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
Text-to-Speech • 2B • Updated • 341k • 433 -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.33M • 2.03k -
Qwen3-TTS Demo
🎙2.29kGenerate speech from text with voice design, cloning, or presets
-
moonshotai/Kimi-K2.5
Image-Text-to-Text • 1T • Updated • 339k • • 2.88k -
ibm-granite/granite-vision-3.3-2b-chart2csv-preview
Image-Text-to-Text • 3B • Updated • 978 • 17 -
Qwen3-ASR Demo
🎙152Transcribe audio to text with timestamps and visualization
-
Qwen/Qwen3-ASR-1.7B
Automatic Speech Recognition • 2B • Updated • 1.33M • • 1.15k -
Qwen/Qwen3.5-35B-A3B
Image-Text-to-Text • 36B • Updated • 1.52M • • 1.52k -
Robot Learning: A Tutorial
Paper • 2510.12403 • Published • 146 -
YatharthS/LuxTTS
Text-to-Speech • Updated • 6.02k • 209 -
LuxTTS
🚀109Space for LuxTTS: a 150x realtime voice cloning TTS model
-
google/gemma-4-31B
Image-Text-to-Text • 33B • Updated • 574k • 550 -
google/gemma-4-31B-it
Image-Text-to-Text • 31B • Updated • 9.59M • • 4.04k -
google/gemma-4-E2B-it
Any-to-Any • 5B • Updated • 3.01M • 1.01k -
google/gemma-4-E4B-it
Any-to-Any • 8B • Updated • 4.42M • 1.65k -
Lightricks/LTX-Video
Image-to-Video • 2B • Updated • 747k • 2.33k -
unsloth/Qwen3.6-27B-GGUF
Image-Text-to-Text • 27B • Updated • 704k • 962 -
Qwen/Qwen3.6-27B
Image-Text-to-Text • 28B • Updated • 2.35M • • 2.32k -
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text • 36B • Updated • 3.52M • • 2.94k -
facebook/sapiens2
Updated • 170 -
google/gemma-4-E4B-it-assistant
Any-to-Any • 78.8M • Updated • 15.4k • 124 -
nvidia/LocateAnything-3B
Image-Text-to-Text • 4B • Updated • 188k • 3.1k -
unsloth/gemma-4-12b-it-GGUF
Image-Text-to-Text • 12B • Updated • 1.6M • 880 -
google/gemma-4-12B-it
Any-to-Any • 12B • Updated • 1.74M • 1.64k -
unsloth/gemma-4-26B-A4B-it-GGUF
Image-Text-to-Text • 25B • Updated • 562k • 1.19k -
MisoLabs/MisoTTS
Text-to-Speech • 8B • Updated • 232 -
unsloth/diffusiongemma-26B-A4B-it-GGUF
Image-Text-to-Text • 25B • Updated • 245k • 408 -
unsloth/Kimi-K2.7-Code
Image-Text-to-Text • 1.1T • Updated • 230 • 21 -
google/diffusiongemma-26B-A4B-it
Image-Text-to-Text • 26B • Updated • 612k • 1.28k -
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
Paper • 2606.19195 • Published • 79 -
LocateAnything
💬504Detect and label objects in images and videos
-
nvidia/Eagle2.5-8B
Image-Text-to-Text • 8B • Updated • 18.3k • 50 -
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
Paper • 2605.27365 • Published • 142 -
HF Realtime Voice
🎙572Realtime voice over WebSocket or WebRTC
-
unsloth/gemma-4-31B-it-NVFP4
Image-Text-to-Text • 23B • Updated • 11.4k • 32 -
unsloth/inkling-GGUF
Image-Text-to-Text • 947B • Updated • 439k • 137 -
lightonai/LightOnOCR-2-1B
Image-Text-to-Text • 1B • Updated • 239k • • 834 -
small-models-for-glam/kraken-ppocrv6-medium
Image-to-Text • Updated • 17 -
Qwen/Qwen-Image-2.1
Text-to-Image • 7B • Updated • 109k • 3.06k -
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 1.29M • • 1.18k -
nvidia/Nemotron-3-Diarization
Voice Activity Detection • 99.2M • Updated • 62.3k • 737 -
FermionResearch/Phonon-2
Automatic Speech Recognition • Updated • 3.75k • 266 -
mohit67890/imajev-4b
Image-Text-to-Text • Updated • 2.11k • 89 -
akhilaaa3/Jev-Omni
Text Classification • 12B • Updated • 4.04k • 380