[FEEDBACK] Inference Providers

#49
by julien-c - opened
Hugging Face org

Any inference provider you love, and that you'd like to be able to access directly from the Hub?

Love that I can call DeepSeek R1 directly from the Hub 🔥

from huggingface_hub import InferenceClient

client = InferenceClient(
    provider="together",
    api_key="xxxxxxxxxxxxxxxxxxxxxxxx"
)

messages = [
    {
        "role": "user",
        "content": "What is the capital of France?"
    }
]

completion = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-R1", 
    messages=messages, 
    max_tokens=500
)

print(completion.choices[0].message)

Is it possible to set a monthly payment budget or rate limits for all the external providers? I don't see such options in billings tab. In case a key is or session token is stolen, it can be quite dangerous to my thin wallet:(

Hugging Face org

@benhaotang you already get spending notifications when crossing important thresholds ($10, $100, $1,000) but we'll add spending limits in the future

@benhaotang you already get spending notifications when crossing important thresholds ($10, $100, $1,000) but we'll add spending limits in the future

Thanks for your quick reply, good to know!

Would be great if you could add Nebius AI Studio to the list :) New inference provider on the market, with the absolute cheapest prices and the highest rate limits...

Could be good to add featherless.ai

TitanML !!

Hi Hugging Face team,

We would like to join Inference Providers as FUSE, a managed inference service hosted in Riyadh, Saudi Arabia.

Provider details
Provider ID: fuse
Hugging Face organization: https://huggingface.co/ai-Humain (Team plan active)
API base URL: https://api.fuse.humainaic.com/v1
API docs: Documentation | HUMAIN FUSE
Contact: Mohammed Alshehri | Maalshehri@humain.com

API compatibility
Standard conversational task through OpenAI-compatible POST /v1/chat/completions, streaming and non-streaming.
Tool calling and structured output (JSON schema) supported.
GET /v1/models returns pricing (USD per million input and output tokens) and context_length for each model.
Billing endpoint for request ID cost lookup in nano-USD: [ready by Mid of October]

Initial models (conversational)
deepseek-ai/DeepSeek-V4.1-Flash
zai-org/GLM-5.3-Flash
Qwen/Qwen3.8-Flash-Next
openai/gpt-oss-120b

Data handling
No routing to third-party providers.
No storage of prompts or outputs after a request completes.
No training on customer content.
We keep only request IDs and costs, for billing.

Why FUSE on the Hub
Hub users in the Middle East get low-latency access to open models with in-Kingdom processing. That matters to Saudi enterprise and government buyers.

Next steps
We are preparing the huggingface.js pull request, and can start model mappings in staging as soon as you enable our organization. Please let us know if you need anything else.

Thank you,
Mohammed | HUMAIN

Hi Hugging Face team,

I'm Jonas, the developer behind Flonno (https://www.flonno.com), based in Switzerland. I'm evaluating a focused self-hosted offering for open-weight models and would like to understand whether it could be a fit for Inference Providers. Our current API forwards requests to external infrastructure; our own GPU deployment is not live yet.

Are you considering new small providers, and are there particular models or regions where additional coverage would be useful? What minimum capacity, performance and availability would you expect for an initial evaluation?

I've read the registration guide and understand that integration includes model mappings, client integration, a billing endpoint and a Team or Enterprise organization plan. Before committing to that work, I'd appreciate guidance on the right onboarding contact and how to validate the fit.

Best regards,
Jonas Boos / Flonno
jonas@flonno.com

Hi Hugging Face team,

I'm José, founder of QDivZero, an AI inference platform built by Valendra Tech in Spain.

We'd love to explore integrating QDivZero with Hugging Face.

QDivZero takes a slightly different approach from traditional inference providers. Instead of maintaining only a fixed catalogue of hosted models, our platform can deploy compatible Hugging Face models on demand with one click.

Given a Hugging Face repository, QDivZero can automatically provision the required GPU infrastructure, select the appropriate inference runtime, launch the model, and expose it through an OpenAI-compatible API.

We currently work with runtimes such as vLLM, SGLang and llama.cpp, and our infrastructure is designed around dynamic GPU provisioning, pay-per-compute usage, and shared inference capacity.

We support two complementary execution models:

  • Dedicated/on-demand capacity, where a model is dynamically deployed for a user or workload.
  • Shared capacity, where multiple users can access already-running model infrastructure, reducing cold starts and improving cost efficiency.

This allows us to dynamically deploy long-tail or custom models while keeping popular models available through shared warm capacity.

We see two particularly interesting integration paths with Hugging Face:

  1. QDivZero as an official Inference Provider for popular models served through our shared/warm capacity.

  2. A "Deploy on QDivZero" workflow that would allow users to launch arbitrary compatible Hugging Face models directly from the Hub.

The second use case is especially interesting to us.

Our goal is to make the experience as simple as:

Hugging Face model → Deploy → GPU provisioned → OpenAI-compatible endpoint

without requiring the user to manually choose GPU types, configure vLLM/SGLang, manage containers, or build inference infrastructure.

For popular models, the experience can instead become:

Hugging Face model → Shared QDivZero capacity → Immediate inference

Over time, models receiving enough demand can move from dedicated/on-demand deployments into shared pools, reducing both startup times and inference costs.

We're based in Europe and are particularly interested in providing a European infrastructure option for open-weight AI models.

Website: https://qdivzero.com

We'd be happy to implement the requirements for Hugging Face Inference Providers and would also love to discuss whether the broader one-click deployment workflow could be interesting for the Hub.

Would this be something the Hugging Face team would be interested in exploring?

Happy to share a demo, architecture details, benchmarks, or jump on a call.

José Carlos García Ortega jose@valendra.tech
Founder — QDivZero / Valendra Tech

Hi HF team — we're TokenWatt (Teampulse Solutions LLC, US), and we'd like to join Inference Providers. Our org: https://huggingface.co/tokenwatt

We serve open-weight LLMs on measured US hardware: every endpoint is sized from our published SLO-goodput benchmarks (https://www.tokenwatt.io/reports.html), with fast 429s instead of queueing. No prompt/completion logging or training (https://www.tokenwatt.io/privacy.html).

Our API is OpenAI-compatible and already implements your contract: base URL https://api.tokenwatt.io/hf, /v1/models with pricing and context length, an Inference-Id header on every response, and a nano-USD billing endpoint.

First models: Qwen/Qwen3-30B-A3B and Qwen/Qwen3-30B-A3B-Instruct-2507 (FP8, accuracy-gated against BF16), with tool calling and structured outputs.

We have the huggingface.js provider helper ready to PR. Could you enable us server-side for the model mapping API? Contact: hello@tokenwatt.io

Hi! We'd like to add impossibl as an inference provider.

impossibl is an OpenAI- and Anthropic-compatible gateway (https://impossibl.com, API at https://api.impossibl.com). For Hugging Face we'd start with conversational models we already serve: openai/gpt-oss-120b, zai-org/GLM-5.3, moonshotai/Kimi-K3, MiniMaxAI/MiniMax-M3, deepseek-ai/DeepSeek-V4-Pro-0813 and Qwen/Qwen3.8-27B.

Where we are:

What we can't do ourselves is get the org enabled for the Model Mapping API and registered server-side. Could someone help with that? Happy to send logos or anything else you need.

Thanks,
Jonathan

Gödel Machines: request to join as an Inference Provider

Hi! We're Gödel Machines (https://huggingface.co/GoedelMachines), an AI lab in Hyderabad, India. We'd like to become an Inference Provider.

  • Endpoint: OpenAI-compatible, live at https://api.goedelmachines.com/v1, with streaming and usage, tool calling and reasoning.
  • First model: Qwen/Qwen3.8-27B (NVIDIA NVFP4 weights), 262k context, on our own RTX PRO 6000 Blackwell GPUs. Median time to first token is about 0.8 s.
  • Differentiator: inference runs in India, with no prompt or completion retention and no training on user data (policy: https://api.goedelmachines.com/legal#privacy).
  • Pricing: $0.10 / 1M input, $0.02 / 1M cached input, $1.90 / 1M output.

We're implementing the Inference-Id header and billing endpoint now, and our org is on the Team plan. We're ready to open the huggingface.js PR. Who should we coordinate with?

Hi Hugging Face team,
I’m following up on our interest in joining Hugging Face Inference Providers. TokenAAS is an OpenAI-compatible API gateway operated by FUELNODE TECHNOLOGIES INC. (Canada).
Our Hugging Face account is weixiaoxin, our organization is https://huggingface.co/tokenaas, our website is https://tokenaas.ai, and our API base URL is https://tokenaas.ai/v1.
We provide unified access to third-party inference services, with API key management, usage metering, and billing; we do not host the model weights ourselves. Could you let us know whether this operating model is eligible, whether you are accepting new provider applications, and where we should submit a formal application? We can provide a provider profile, technical details, supporting documents, and evaluation credentials through a private channel.
I have also contacted your team by email. You can reach us at admin@tokenaas.ai.
Thank you,
The TokenAAS team

Hi Hugging Face team,

We’re preparing a self-hosted inference service and would like to explore joining Hugging Face Inference Providers.

Our initial model is Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8. We have implemented and locally tested:

  • An OpenAI-compatible chat completions API
  • Streaming and non-streaming responses
  • Tool calling and structured outputs
  • A /v1/models endpoint exposing pricing and context length
  • Unique Inference-Id response headers
  • A persistent billing endpoint returning per-request costs in nano-USD, with idempotent lookups

Public HTTPS deployment and sustained reliability testing are currently in progress.

Are you currently accepting new providers? Could you advise on the eligibility requirements, the appropriate onboarding contact, and the next steps for integration review and server-side registration?

We’d be happy to share our API documentation, test results, and evaluation credentials through a private channel.

Contact: blake19950908@gmail.com

Thank you!

Sign up or log in to comment