together_ai

NVIDIA: Nemotron-3 Ultra 550B A55B

Unmapped
Context

1M

Input

text

Output

text

Tool calling

Supported

About this model

NVIDIA Nemotron-3 Ultra 550B A55B is a frontier-scale large language model from NVIDIA, designed for strong agentic, reasoning, and conversational capabilities. It employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2, MoE, and Attention layers, alongside Multi-Token Prediction (MTP) for faster inference. With a 1-million-token context window and a configurable reasoning mode, it excels at complex autonomous agents, long-context analysis, and deep research workflows.

Input modalities

text

Accepted as model input

Providers

Available routing options for this model through OneInfer.

Available

Input tokens

$0.600 / 1M tokens

Output tokens

$3.600 / 1M tokens

Routing

OneInfer optimized

Pricing

Current OneInfer pricing for this model.

UsagePrice
Input tokens$0.600 / 1M tokens
Output tokens$3.600 / 1M tokens
Cached input tokens$0.200 / 1M tokens

Performance

Published evaluation results associated with this model.

Reasoning

GPQA (no tools)87
MMLU-Pro86.8

Coding

SWE-Bench Verified71.9
LiveCodeBench (v6)89

Agentic

Terminal Bench 2.167.2
TauBench V3 (Telecom)98.3

Long Context

RULER (1M)94.7

API example

curl https://api.oneinfer.ai/v1/ula/chat/completions \
  -H "Authorization: Bearer $JWT_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Unmapped",
    "messages": [
      { "role": "user", "content": "Hello!" }
    ]
  }'