- Debian Inference Portal
- What this service provides
- Available models and providers
- Quick start guide
- Getting a key
- Configuring tools
- FAQ
- Who can use this service?
- I got "Salsa account not linked" — what now?
- What can I use it for?
- How much does it cost me?
- What happens if I hit my budget?
- Can I renew my credits?
- Which model should I use?
- I lost my key.
- My key stopped working.
- Can I use this from any OpenAI-compatible tool?
- Who owns the output, and does Scaleway train on my data?
- Support
What this service provides
The Debian Inference Portal is a self-service way for Debian contributors to use shared LLM inference. You log in with your Salsa account, create API keys, and track your spend — all through your browser. The portal provisions you on a shared, OpenAI-compatible inference proxy, so any tool that speaks the OpenAI API can use your key.
It is a shared, fair-use service funded by a fixed monthly pool of credits. Your weekly budget is a soft limit, in place for safety reasons: it paces usage so the shared pool lasts for everyone, and you can renew it from the dashboard once it is spent. Please still avoid long or very heavy workloads.
Available models and providers
Scaleway is a European cloud provider owned by the
Iliad Group (the parent of Free),
delivering public cloud, bare metal, AI and managed services. Among those services is
Generative APIs, a serverless,
OpenAI-compatible API for serving LLMs — you pay per token and don't manage any
infrastructure. The models below are served through it: requests are proxied to
api.scaleway.ai, and you don't need your own Scaleway account — the service holds the
provider credentials and forwards your calls.
This service is possible thanks to Scaleway, which kindly sponsors the Debian project with a monthly allocation of credits to cover the inference costs for contributors.
Model ID (use this in model) |
Model |
|---|---|
scaleway/glm-5.2 |
Zhipu AI's GLM-5.2, a general-purpose model |
scaleway/deepseek-v4-flash-0731 |
DeepSeek's V4 Flash, a fast, lightweight model |
scaleway/qwen3.8-27b |
Qwen's 3.8-27B, a general-purpose model |
Both models are in preview at Scaleway. They are billed by token, but note that GLM-5.2 currently does not differentiate cached (input) tokens, so re-reading long context is charged at the full rate and can get comparatively expensive. Re-reading long context is typical in the context of agentic software development (agents re-read the accumulated conversation on every step), so this adds up quickly. DeepSeek V4 Flash is recommended for most uses.
The list above is not exhaustive — run the models endpoint with your key to see exactly what is currently exposed:
curl -s https://inference.debian.net/v1/models -H "Authorization: Bearer $DEBIAN_INFERENCE_KEY"
Quick start guide
- Sign in — you'll land on your dashboard, where you can create a key.
- Create a key from your dashboard and copy the value shown. It is only displayed once.
- Choose a coding agent — OpenCode is the recommended choice. None of the modern coding agents are packaged in Debian, unfortunately.
- Isolate your coding agent — a coding agent reads your files and runs commands, so treat
it like any untrusted program: it could leak your private keys or escalate to root. Run it
sandboxed, for example with
bubblewrap, which is packaged in Debian. A wrapper for OpenCode looks like this:
bwrap \
--ro-bind /usr /usr \
--symlink usr/bin /bin \
--symlink usr/lib /lib \
--symlink usr/lib64 /lib64 \
--ro-bind /etc/ssl /etc/ssl \
--ro-bind /etc/alternatives /etc/alternatives \
--ro-bind /etc/dpkg/origins /etc/dpkg/origins \
--ro-bind /var/lib/dpkg/status /var/lib/dpkg/status \
--ro-bind /etc/resolv.conf /etc/resolv.conf \
--ro-bind $HOME/.gitconfig $HOME/.gitconfig \
--bind $HOME/.config/opencode $HOME/.config/opencode \
--bind $HOME/.cache/opencode $HOME/.cache/opencode \
--bind $HOME/.opencode $HOME/.opencode \
--bind $HOME/.local/share/opencode $HOME/.local/share/opencode \
--bind $HOME/.local/state/opencode $HOME/.local/state/opencode \
--proc /proc \
--dev /dev \
--tmpfs /tmp \
--bind $(pwd) $(pwd) \
--chdir $(pwd) \
--unshare-all \
--share-net \
--die-with-parent \
/path/to/opencode "$@"
This is not a full containment: the sandboxed process still has the terminal and can reach local sockets that don't go through the filesystem (abstract Unix sockets, TCP sockets, and so on).
- Configure your coding agent as described below.
Getting a key
- Sign in with Salsa (top right). Only Debian Developers and Debian
Maintainers are granted access; the portal checks your status against
nm.debian.org. - You land on your dashboard, which shows your budget and the keys you've created.
- Create a key, copy the generated value, and use it as a Bearer token against the proxy.
Things to know about keys:
- Keys are created per-user and inherit your budget and the models the proxy exposes.
- A new key expires after 90 days by default. Use Renew on the dashboard to extend one by another 90 days. (This extends the key's lifetime; it is different from Renew credits, which resets your spending — see the FAQ.)
- A key value is shown to you only once, right after creation. The portal never stores the raw value — if you lose it, revoke it and create a new one.
- You can revoke a key at any time; anything using it will stop working immediately.
The proxy speaks the OpenAI API: the classic
Chat Completions interface (POST /v1/chat/completions with a messages payload) and the
newer Responses API (POST /v1/responses). The examples below use Chat Completions, which
is the most widely supported. Use your key as a Bearer token and set model to one of the
exposed models (e.g. scaleway/deepseek-v4-flash-0731).
curl
export DEBIAN_INFERENCE_KEY="sk-your-key"
# List the models available to you
curl -s https://inference.debian.net/v1/models \
-H "Authorization: Bearer $DEBIAN_INFERENCE_KEY"
# A simple chat completion
curl -s https://inference.debian.net/v1/chat/completions \
-H "Authorization: Bearer $DEBIAN_INFERENCE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "scaleway/deepseek-v4-flash-0731",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Python (requests)
import os
import requests
key = os.environ["DEBIAN_INFERENCE_KEY"]
resp = requests.post(
"https://inference.debian.net/v1/chat/completions",
headers={"Authorization": f"Bearer {key}"},
json={
"model": "scaleway/deepseek-v4-flash-0731",
"messages": [{"role": "user", "content": "Hello!"}],
},
)
resp.raise_for_status()
print(resp.json()["choices"][0]["message"]["content"])
OpenCode
Add a custom OpenAI-compatible provider to your opencode.json (project-level, or the global
~/.config/opencode/opencode.json):
{
"model": "debian-inference/scaleway/deepseek-v4-flash-0731",
"provider": {
"debian-inference": {
"npm": "@ai-sdk/openai-compatible",
"name": "Debian Inference",
"api": "https://inference.debian.net/v1",
"env": [
"DEBIAN_INFERENCE_KEY"
],
"models": {
"scaleway/deepseek-v4-flash-0731": {
"name": "DeepSeek V4 Flash 0731",
"family": "deepseek-flash",
"release_date": "2026-07-31",
"attachment": false,
"reasoning": true,
"temperature": true,
"tool_call": true,
"cost": {
"input": 0.4,
"output": 0.8,
"cache_read": 0.08
},
"limit": {
"context": 256000,
"output": 16384
},
"modalities": {
"input": ["text"],
"output": ["text"]
},
"interleaved": true
},
"scaleway/glm-5.2": {
"name": "GLM-5.2",
"family": "glm",
"release_date": "2026-06-13",
"attachment": false,
"reasoning": true,
"temperature": true,
"tool_call": true,
"cost": {
"input": 1.8,
"output": 5.5
},
"limit": {
"context": 256000,
"output": 16384
},
"modalities": {
"input": ["text"],
"output": ["text"]
}
},
"scaleway/qwen3.8-27b": {
"name": "Qwen3.8-27B",
"family": "qwen",
"release_date": "2026-08-01",
"attachment": false,
"reasoning": true,
"temperature": true,
"tool_call": true,
"cost": {
"input": 0.6,
"output": 3.3,
"cache_read": 0.12
},
"limit": {
"context": 256000,
"output": 32768
},
"modalities": {
"input": ["text"],
"output": ["text"]
},
"interleaved": true
}
}
}
}
}
Then start OpenCode, use /connect, search for Debian Inference, and enter your API key. The example config above sets DeepSeek V4 Flash 0731 as the default model — use /models to switch to GLM-5.2 if you prefer it. Then make your first prompt.
Pi (pi.dev)
Pi is a minimal, extensible agent harness. Add a custom
OpenAI-compatible provider to your ~/.pi/agent/models.json to use the service:
{
"providers": {
"debian-inference": {
"baseUrl": "https://inference.debian.net/v1",
"api": "openai-completions",
"apiKey": "$DEBIAN_INFERENCE_KEY",
"models": [
{
"id": "scaleway/deepseek-v4-flash-0731",
"name": "DeepSeek V4 Flash 0731",
"reasoning": true,
"contextWindow": 256000,
"maxTokens": 16384,
"cost": {
"input": 0.4,
"output": 0.8,
"cacheRead": 0.08,
"cacheWrite": 0.4
}
},
{
"id": "scaleway/glm-5.2",
"name": "GLM-5.2",
"reasoning": true,
"contextWindow": 256000,
"maxTokens": 16384,
"cost": {
"input": 1.8,
"output": 5.5,
"cacheRead": 1.8,
"cacheWrite": 1.8
}
},
{
"id": "scaleway/qwen3.8-27b",
"name": "Qwen3.8-27B",
"reasoning": true,
"contextWindow": 256000,
"maxTokens": 32768,
"cost": {
"input": 0.6,
"output": 3.3,
"cacheRead": 0.12,
"cacheWrite": 0.6
}
}
]
}
}
}
The file reloads each time you open /model. Set DEBIAN_INFERENCE_KEY to your
key (for example via /login for the provider, or exporting the env var) and
pick a model with /model.
DebGPT
DebGPT is a
Debian-flavored terminal LLM tool, packaged in Debian. Use the version in
testing/forky — it installs cleanly on stable too — because it is the one with
tool support (calling of external programs). Point it
at the portal with the openai frontend in ~/.config/debgpt/config.toml:
frontend = "openai"
openai_base_url = "https://inference.debian.net/v1"
openai_model = "scaleway/deepseek-v4-flash-0731"
openai_api_key = "sk-your-key"
FAQ
Who can use this service?
Active Debian Developers and Debian Maintainers. Access is granted on sign-in, based on your
status on nm.debian.org.
I got "Salsa account not linked" — what now?
The portal resolves your Debian identity from your Salsa account through nm.debian.org. If
your Salsa login is not yet linked to your Debian identity there, log into
nm.debian.org with both your Salsa account and your Debian SSO
account — that establishes the link — then sign in here again.
What can I use it for?
Use it for work that benefits the Debian project — for example packaging, bug triage and fixes, tooling, documentation, or anything else you're doing as a contributor. It isn't for unrelated personal tasks. The underlying resources are provided free of charge by sponsors, so please be prepared to describe, in a short report, what you used it for if asked — such reports help make the case to sponsors for continuing and growing the service.
How much does it cost me?
Nothing directly — the service is provided for the Debian project and funded by a fixed monthly pool of credits from its sponsors. Spending is tracked per user over a weekly window that resets automatically, with defaults tiered by status:
| Status | Default weekly budget |
|---|---|
| Debian Developer (DD) | $25 |
| Debian Maintainer (DM) | $10 |
This weekly budget is a soft limit, in place for safety reasons — not a hard cap on what you may use. It exists to pace usage so the shared monthly pool isn't drained by a few heavy users, which would leave nothing for everyone else. The limits are somewhat arbitrary and intended as a starting point; they may evolve over time as we learn how the service is used. Your current spend and the next automatic reset time are shown on the dashboard.
What happens if I hit my budget?
Budgets are enforced by the proxy at the user level, so the total spend across all of your keys counts toward your limit. Once you reach it, new requests fail until you renew your credits (below) or the weekly window resets. Hitting the limit is normal for intensive work — it isn't a sign that you've done something wrong.
Can I renew my credits?
Yes — renewing is the normal way to keep working once you've spent your weekly credits. The Renew credits button on the dashboard resets your weekly spend to 0, giving you a fresh window without changing your allowance. You can renew as soon as your credits are spent, subject to a safety check: because everyone shares one monthly pool, a renewal is only allowed while the predicted organization-wide monthly spend still fits comfortably inside the pool (with a safety margin). The dashboard shows whether a renewal is currently available and the predicted monthly spend; when it is, use the Renew credits button. If it isn't available, wait for the weekly reset or contact the team.
Which model should I use?
scaleway/deepseek-v4-flash-0731 is the recommended default: it's fast, lightweight, and
cheaper than GLM-5.2 because it differentiates cached tokens. scaleway/glm-5.2 is a strong
general-purpose model, but it's in preview and currently charges cached input at the full
rate. Use /v1/models to see what's currently available.
I lost my key.
Keys are shown only once at creation. Revoke it on the dashboard and create a new one.
My key stopped working.
Check the expiry date on the dashboard and renew if needed. If you've exceeded your budget, renew your credits from the dashboard, or wait for the weekly reset.
Can I use this from any OpenAI-compatible tool?
Yes — any tool that lets you set a custom base URL (https://inference.debian.net/v1) and a
Bearer token can use your key (see the examples above).
Who owns the output, and does Scaleway train on my data?
Outputs generated by the service belong to you. Scaleway does not claim any copyright or ownership over generated content.
Scaleway applies a Zero Data Retention Policy: prompt content is not stored, and is explicitly not used for training, retraining, or improving the base models. Data is not accessible to the LLM creators or to third parties. Only anonymized metadata (token counts, HTTP status codes) is retained for up to 6 months for performance monitoring. Data is hosted in Paris, France, and as a European company Scaleway is not subject to extraterritorial laws such as the US Cloud Act. See Scaleway's Generative APIs privacy policy and Specific Conditions for AI Services for the full details.
Support
Questions, feedback, or problems? Contact the Debian AI team by filing an issue at inference-support.
Inference resources sponsored by Scaleway