Nemotron 3.5 Lightning 30B API pricing and speed | RunInfra

RunInfra

2 min read Original article ↗

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Nemotron 3.5 Lightning 30B is an LLM listed in RunInfra Model APIs. RunInfra serves it as nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 at $0.05 per 1M input tokens and $0.15 per 1M output tokens. Its context window is 262,144 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.

Pricing

USD, pay per token

per 1M input tokens
$0.05

per 1M cached input tokens
$0.01

per 1M output tokens
$0.15

Measured performance

Output speed

540.9output tokens per second, model only

Time to first token

67milliseconds to first reasoning token

Access

Confirm how your client reaches this model.

Provider
NVIDIA

API compatibility
OpenAI-compatible chat completions
Anthropic compatibility
Anthropic-compatible Messages, POST /v1/messages

Capacity

Check the limits your workload must fit.

Context window
262,144 tokens

Maximum request size
3.5 MB per request
Maximum generated output
262,144 tokens

Capabilities

See which request modes the API supports.

Tool calling
Supported

Precision
BF16, unquantized
View full spec

Gateway compatibility
OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways

Prefix caching
Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.

Cache retention
Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.

Upstream model
View model