Early access · Paid API credit is live. New accounts still get free starter credit.
PricingDocsModelsFAQEarnBlogAboutSign in →

Qwen/Qwen3-8B-GGUF

Qwen3 8B Q4_K_M API

Yes. Umbra lists this public Hugging Face GGUF in its hosted API catalog. When an eligible attested provider is online, call it through the OpenAI- or Anthropic-compatible API using the exact model ID below.

Model IDhf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m
OpenAI base URLhttps://api.tryumbra.dev/v1
Catalog pricing$0.10/M input · $0.30/M output

The exact public artifact

These fields form the complete pinned artifact descriptor providers are approved to serve.

Hugging Face repository
Qwen/Qwen3-8B-GGUF
GGUF file
Qwen3-8B-Q4_K_M.gguf
Revision
7c41481f57cb95916b40956ab2f0b139b296d974
SHA-256
d98cdcbd03e17ce47681435b5150e34c1417f50b5c0019dd560e4882c5745785
Architecture
qwen3
Quantization
Q4_K_M
License
apache-2.0
Minimum unified memory
8 GB
Context window
40,960 tokens
Reasoning
Yes, with skip-thinking support

Call it with the SDK you already use

The API key is the only secret. Never place a real key in a model card, browser page, or repository.

OpenAI Python

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tryumbra.dev/v1",
    api_key="umbra-...",
)

response = client.chat.completions.create(
    model="hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m",
    messages=[{"role": "user", "content": "Hello"}],
)

print(response.choices[0].message.content)

OpenAI Node.js

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.tryumbra.dev/v1",
  apiKey: "umbra-...",
});

const response = await client.chat.completions.create({
  model: "hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m",
  messages: [{ role: "user", content: "Hello" }],
});

console.log(response.choices[0].message.content);

Anthropic Python

from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.tryumbra.dev",
    api_key="umbra-...",
    default_headers={"anthropic-version": "2023-06-01"},
)

message = client.messages.create(
    model="hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello"}],
)

print(message.content[0].text)

For this model’s Hugging Face card

If you maintain the repository, this short block gives users an honest hosted-API link. Keep it only while the catalog page remains accurate.

Hosted API block

## Hosted API

This GGUF is available through [Umbra](https://tryumbra.dev/models/hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m/), an OpenAI- and Anthropic-compatible API for public Hugging Face models served from attested Apple Silicon.

- Model ID: `hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m`
- API details and live availability: https://tryumbra.dev/models/hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m/

Questions about the Qwen3 8B API

Is there an API for Qwen/Qwen3-8B-GGUF?

Yes. Umbra lists this public Hugging Face GGUF under the model ID hf-qwen-qwen3-8b-gguf-qwen3-8b-q4-k-m. Calls are accepted while an eligible attested provider is online.

Is the Qwen3 8B API OpenAI-compatible?

Yes. Use https://api.tryumbra.dev/v1 as the OpenAI base URL. Umbra also exposes an Anthropic-compatible messages endpoint.

Which Qwen3 8B artifact does Umbra serve?

The catalog pins Qwen/Qwen3-8B-GGUF · 7c41481f57cb95916b40956ab2f0b139b296d974 · Qwen3-8B-Q4_K_M.gguf with GGUF SHA-256 d98cdcbd03e17ce47681435b5150e34c1417f50b5c0019dd560e4882c5745785.

Does Umbra retain prompts sent to Qwen3 8B?

No. Prompts and outputs are decrypted in process, held in memory for the request, and never logged or persisted.

How much Apple unified memory is needed to host Qwen3 8B?

The current catalog memory floor is 8 GB of unified memory.