nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Nemotron 3.5 Lightning 30B is an LLM listed in RunInfra Model APIs. RunInfra serves it as nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 at $0.05 per 1M input tokens and $0.15 per 1M output tokens. Its context window is 262,144 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.
Pricing
USD, pay per token
- per 1M input tokens
- $0.05
- per 1M cached input tokens
- $0.01
- per 1M output tokens
- $0.15
Measured performance
Output speed
540.9output tokens per second, model only
Time to first token
67milliseconds to first reasoning token
Access
Confirm how your client reaches this model.
- Provider
- NVIDIA
- API compatibility
- OpenAI-compatible chat completions
- Anthropic compatibility
- Anthropic-compatible Messages, POST /v1/messages
Capacity
Check the limits your workload must fit.
- Context window
- 262,144 tokens
- Maximum request size
- 3.5 MB per request
- Maximum generated output
- 262,144 tokens
Capabilities
See which request modes the API supports.
- Tool calling
- Supported
- Precision
- BF16, unquantized
View full spec
- Gateway compatibility
- OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways
- Prefix caching
- Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
- Cache retention
- Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
- Upstream model
- View model