Early access · Paid API credit is live. New accounts still get free starter credit.
PricingDocsModelsFAQEarnBlogAboutSign in →

yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF

gemma 4 12B API

Yes. Umbra lists this public Hugging Face GGUF in its hosted API catalog. When an eligible attested provider is online, call it through the OpenAI- or Anthropic-compatible API using the exact model ID below.

Model IDhf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m
OpenAI base URLhttps://api.tryumbra.dev/v1
Catalog pricing$0.20/M input · $0.60/M output

The exact public artifact

These fields come from Umbra’s verified catalog record. They identify what providers are approved to serve.

Hugging Face repository
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
GGUF file
gemma4-coding-Q4_K_M.gguf
Revision
1380be1796e559fca96b4107599285cab3ddbb92
SHA-256
1fe90b72e105d7bc71650aa59883edece3e84751af489075217a7ae717b1fe8d
Architecture
gemma4
Quantization
Q4_K_M
License
apache-2.0
Minimum unified memory
9 GB
Context window
262,144 tokens
Reasoning
Yes, with skip-thinking support

Call it with the SDK you already use

The API key is the only secret. Never place a real key in a model card, browser page, or repository.

OpenAI Python

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tryumbra.dev/v1",
    api_key="umbra-...",
)

response = client.chat.completions.create(
    model="hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m",
    messages=[{"role": "user", "content": "Hello"}],
)

print(response.choices[0].message.content)

OpenAI Node.js

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.tryumbra.dev/v1",
  apiKey: "umbra-...",
});

const response = await client.chat.completions.create({
  model: "hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m",
  messages: [{ role: "user", content: "Hello" }],
});

console.log(response.choices[0].message.content);

Anthropic Python

from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.tryumbra.dev",
    api_key="umbra-...",
    default_headers={"anthropic-version": "2023-06-01"},
)

message = client.messages.create(
    model="hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello"}],
)

print(message.content[0].text)

For this model’s Hugging Face card

If you maintain the repository, this short block gives users an honest hosted-API link. Keep it only while the catalog page remains accurate.

Hosted API block

## Hosted API

This GGUF is available through [Umbra](https://tryumbra.dev/models/hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m/), an OpenAI- and Anthropic-compatible API for public Hugging Face models served from attested Apple Silicon.

- Model ID: `hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m`
- API details and live availability: https://tryumbra.dev/models/hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m/

Questions about the gemma 4 12B coder fable5 composer2.5 v1 API

Is there an API for yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF?

Yes. Umbra lists this public Hugging Face GGUF under the model ID hf-yuxinlu1-gemma-4-12b-coder-fable5-composer2.5-v1-gguf-gemma4-coding-q4-k-m. Calls are accepted while an eligible attested provider is online.

Is the gemma 4 12B coder fable5 composer2.5 v1 API OpenAI-compatible?

Yes. Use https://api.tryumbra.dev/v1 as the OpenAI base URL. Umbra also exposes an Anthropic-compatible messages endpoint.

Which gemma 4 12B coder fable5 composer2.5 v1 artifact does Umbra serve?

The catalog pins yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF · 1380be1796e559fca96b4107599285cab3ddbb92 · gemma4-coding-Q4_K_M.gguf with GGUF SHA-256 1fe90b72e105d7bc71650aa59883edece3e84751af489075217a7ae717b1fe8d.

Does Umbra retain prompts sent to gemma 4 12B coder fable5 composer2.5 v1?

No. Prompts and outputs are decrypted in process, held in memory for the request, and never logged or persisted.

How much Apple unified memory is needed to host gemma 4 12B coder fable5 composer2.5 v1?

The current catalog memory floor is 9 GB of unified memory.