256K
text
text
Supported
About this model
NVIDIA Nemotron-3 Nano 30B A3B is NVIDIA's open reasoning model featuring a hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture. It consists of 23 Mamba-2 and MoE layers, plus 6 Attention layers. The model is built for building specialized AI agents, chatbots, and RAG systems, providing high compute efficiency. It natively supports configurable reasoning depth through a 'thinking budget' and acts as a general-purpose reasoning and chat model intended for English, coding languages, and several other supported languages.
Input modalities
text
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Input tokens
$0.050 / 1M tokens
Output tokens
$0.200 / 1M tokens
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Input tokens | $0.050 / 1M tokens |
| Output tokens | $0.200 / 1M tokens |
Performance
Published evaluation results associated with this model.
General Knowledge
Reasoning
Agentic
API example
curl https://api.oneinfer.ai/v1/ula/chat/completions \
-H "Authorization: Bearer $JWT_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B",
"messages": [
{ "role": "user", "content": "Hello!" }
]
}'