Image-Text-to-Text
PaddleOCR
Safetensors
MLX
English
Chinese
multilingual
paddleocr_vl
ERNIE4.5
PaddlePaddle
image-to-text
ocr
document-parse
layout
table
formula
chart
seal
spotting
conversational
custom_code
4-bit precision
Instructions to use mlx-community/PaddleOCR-VL-1.5-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PaddleOCR
How to use mlx-community/PaddleOCR-VL-1.5-4bit with PaddleOCR:
# See https://www.paddleocr.ai/latest/version3.x/pipeline_usage/PaddleOCR-VL.html to installation from paddleocr import PaddleOCRVL pipeline = PaddleOCRVL(pipeline_version="mlx-community/PaddleOCR-VL-1.5-4bit") output = pipeline.predict("path/to/document_image.png") for res in output: res.print() res.save_to_json(save_path="output") res.save_to_markdown(save_path="output") - MLX
How to use mlx-community/PaddleOCR-VL-1.5-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mlx-community/PaddleOCR-VL-1.5-4bit") config = load_config("mlx-community/PaddleOCR-VL-1.5-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Update README.md
#1
by umtksa - opened
README.md
CHANGED
|
@@ -35,3 +35,48 @@ pip install -U mlx-vlm
|
|
| 35 |
```bash
|
| 36 |
python -m mlx_vlm.generate --model mlx-community/PaddleOCR-VL-1.5-4bit --max-tokens 100 --temperature 0.0 --prompt "Describe this image." --image <path_to_image>
|
| 37 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
```bash
|
| 36 |
python -m mlx_vlm.generate --model mlx-community/PaddleOCR-VL-1.5-4bit --max-tokens 100 --temperature 0.0 --prompt "Describe this image." --image <path_to_image>
|
| 37 |
```
|
| 38 |
+
|
| 39 |
+
```python
|
| 40 |
+
import sys
|
| 41 |
+
import warnings
|
| 42 |
+
|
| 43 |
+
from PIL import Image
|
| 44 |
+
from mlx_vlm import load, generate
|
| 45 |
+
from mlx_vlm.prompt_utils import apply_chat_template
|
| 46 |
+
from mlx_vlm.utils import load_config
|
| 47 |
+
|
| 48 |
+
image_path = sys.argv[1]
|
| 49 |
+
model_path = "mlx-community/PaddleOCR-VL-1.5-4bit"
|
| 50 |
+
|
| 51 |
+
model, processor = load(model_path)
|
| 52 |
+
config = load_config(model_path)
|
| 53 |
+
|
| 54 |
+
image = [Image.open(image_path)]
|
| 55 |
+
|
| 56 |
+
# PaddleOCR-VL task-specific prompt'lar:
|
| 57 |
+
# "OCR:" → general OCR jobs
|
| 58 |
+
# "Table Recognition:" → table
|
| 59 |
+
# "Formula Recognition:" → formulas
|
| 60 |
+
# "Chart Recognition:" → graphics
|
| 61 |
+
prompt = "OCR:"
|
| 62 |
+
|
| 63 |
+
formatted_prompt = apply_chat_template(
|
| 64 |
+
processor,
|
| 65 |
+
config,
|
| 66 |
+
prompt,
|
| 67 |
+
num_images=len(image)
|
| 68 |
+
)
|
| 69 |
+
|
| 70 |
+
result = generate(
|
| 71 |
+
model,
|
| 72 |
+
processor,
|
| 73 |
+
formatted_prompt,
|
| 74 |
+
image,
|
| 75 |
+
max_tokens=1024,
|
| 76 |
+
temperature=0.0,
|
| 77 |
+
verbose=False,
|
| 78 |
+
)
|
| 79 |
+
|
| 80 |
+
output = result if isinstance(result, str) else result.text
|
| 81 |
+
print(output)
|
| 82 |
+
```
|