Early access · Paid API credit is live. New accounts still get free starter credit.
PricingDocsModelsFAQEarnBlogAboutSign in →

Qwen/Qwen2.5-0.5B-Instruct-GGUF

Qwen2.5 0.5B Instruct F16 API

Yes. Umbra lists this public Hugging Face GGUF in its hosted API catalog. When an eligible attested provider is online, call it through the OpenAI- or Anthropic-compatible API using the exact model ID below.

Model IDhf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16
OpenAI base URLhttps://api.tryumbra.dev/v1
Catalog pricing$0.03/M input · $0.10/M output

Published artifact provenance

This catalog entry is missing at least one field needed for an exact artifact claim. Verify the repository, revision, filename, and digest before relying on it.

Hugging Face repository
Qwen/Qwen2.5-0.5B-Instruct-GGUF
GGUF file
Not published
Revision
9217f5db79a29953eb74d5343926648285ec7e67
SHA-256
8e0ae26000627ed62de0e78e41860af70094558b9d2913385c842a6aa06cf3fc
Architecture
qwen2
Quantization
F16
License
apache-2.0
Minimum unified memory
8 GB
Context window
8,192 tokens
Reasoning
Not asserted by the catalog

Call it with the SDK you already use

The API key is the only secret. Never place a real key in a model card, browser page, or repository.

OpenAI Python

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tryumbra.dev/v1",
    api_key="umbra-...",
)

response = client.chat.completions.create(
    model="hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16",
    messages=[{"role": "user", "content": "Hello"}],
)

print(response.choices[0].message.content)

OpenAI Node.js

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.tryumbra.dev/v1",
  apiKey: "umbra-...",
});

const response = await client.chat.completions.create({
  model: "hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16",
  messages: [{ role: "user", content: "Hello" }],
});

console.log(response.choices[0].message.content);

Anthropic Python

from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.tryumbra.dev",
    api_key="umbra-...",
    default_headers={"anthropic-version": "2023-06-01"},
)

message = client.messages.create(
    model="hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello"}],
)

print(message.content[0].text)

For this model’s Hugging Face card

If you maintain the repository, this short block gives users an honest hosted-API link. Keep it only while the catalog page remains accurate.

Hosted API block

## Hosted API

This GGUF is available through [Umbra](https://tryumbra.dev/models/hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16/), an OpenAI- and Anthropic-compatible API for public Hugging Face models served from attested Apple Silicon.

- Model ID: `hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16`
- API details and live availability: https://tryumbra.dev/models/hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16/

Questions about the Qwen2.5 0.5B Instruct API

Is there an API for Qwen/Qwen2.5-0.5B-Instruct-GGUF?

Yes. Umbra lists this public Hugging Face GGUF under the model ID hf-qwen-qwen2.5-0.5b-instruct-gguf-qwen2-5-0-5b-instruct-fp16. Calls are accepted while an eligible attested provider is online.

Is the Qwen2.5 0.5B Instruct API OpenAI-compatible?

Yes. Use https://api.tryumbra.dev/v1 as the OpenAI base URL. Umbra also exposes an Anthropic-compatible messages endpoint.

Which Qwen2.5 0.5B Instruct artifact does Umbra serve?

Umbra’s catalog record does not yet publish every field needed to identify one exact artifact: repository, immutable revision, GGUF filename, and SHA-256 digest. Review the provenance fields below before use.

Does Umbra retain prompts sent to Qwen2.5 0.5B Instruct?

No. Prompts and outputs are decrypted in process, held in memory for the request, and never logged or persisted.

How much Apple unified memory is needed to host Qwen2.5 0.5B Instruct?

The current catalog memory floor is 8 GB of unified memory.