Automatic Speech Recognition
Core ML
NeMo
English
coremltools
asr
speech
parakeet
nvidia
int8-per-channel-symmetric
Instructions to use OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8: direct link, hf CLI and curl.
- Browser
- Download file 1.31 kB
-
https://huggingface.co/OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8/resolve/main/README.md
- Command line
-
hf download hf://OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8/README.md
-
curl -L -o README.md https://huggingface.co/OpenVoiceOS/parakeet-ctc-0.6b-vi-coreml-int8/resolve/main/README.md
1.31 kB
metadata
license: cc-by-4.0
language:
- en
tags:
- asr
- speech
- coreml
- nemo
- parakeet
- nvidia
- int8-per-channel-symmetric
library_name: coremltools
pipeline_tag: automatic-speech-recognition
base_model: nvidia/parakeet-ctc-0.6b-Vietnamese
parakeet-ctc-0.6b-vi-coreml-int8
CoreML conversion of nvidia/parakeet-ctc-0.6b-Vietnamese — INT8 PER CHANNEL SYMMETRIC quantized.
| Architecture | CTC |
| Language | English |
| Sample rate | 16000 Hz |
| Max audio | 15.0s |
| Vocab size | 1024 |
| Framework | NVIDIA NeMo → CoreML (coremltools) |
Components
| File | Component | Best compute |
|---|---|---|
parakeet_mel_encoder.mlpackage |
mel_encoder | ANE / GPU |
parakeet_ctc_decoder.mlpackage |
ctc_decoder | ANE / GPU |
Usage
pip install ovos-stt-plugin-coreml
from ovos_stt_plugin_coreml import CoremlSTT
from ovos_plugin_manager.utils.audio import AudioFile
stt = CoremlSTT(config={"metadata": "metadata.json"})
with AudioFile("speech.wav") as f:
audio = f.read()
print(stt.execute(audio))