GPU training

Run GPU training jobs from the OneInfer console.

Upload a Python training script or configure a job through guided fields. Choose the GPU resources, submit the workload, and track its lifecycle without leaving the console.

Configure

Script or guided parameters

Use your own training code or fill in model, dataset, and training settings.

Allocate

Job-specific GPU resources

Set the GPU type and count, disk capacity, spot usage, and provider preference.

Operate

Visible job lifecycle

Review job status and resources, refresh results, or cancel active work.

What the platform supports

The controls needed to submit and manage training work.

The training page is connected to the same GPU, storage, and job-management workflows already available in OneInfer.

Use a Python training script

Upload your own Python script when you need complete control over the training code and runtime arguments.

Configure model and dataset inputs

Use the guided form to provide model and dataset paths, training parameters, and optional LoRA settings.

Choose the compute for each job

Select the GPU model and quantity, disk size, spot preference, and preferred cloud capacity for the workload.

Track the job lifecycle

See submitted jobs, refresh their latest status, review assigned resources, and cancel queued or running work.

Keep data and artifacts available

Use OneInfer storage for persistent datasets, checkpoints, and model artifacts that need to survive beyond a job.

Request clustered capacity

For workloads that need coordinated GPU capacity, submit a cluster request from the GPU marketplace workflow.

Training workflow

From configuration to a running batch job.

Each job keeps its training configuration and compute selection together, so the console can show what was requested and what is running.

01

Choose how to configure the job

Start with an uploaded Python script or use the guided parameter form.

02

Add training inputs

Provide the model, dataset, hyperparameters, and optional LoRA configuration.

03

Select GPU resources

Choose the GPU type, count, storage, spot preference, and cloud preference.

04

Submit and monitor

Follow the job status from the console and cancel queued or running work when needed.

Connected workflows

Keep training data and deployment work in one platform.

Use persistent storage for datasets and model artifacts. When a supported custom model is ready, continue through the dedicated endpoint workflow to prepare it for inference.

Submit your first GPU training job.

Open the Training console to configure resources, submit a batch job, and monitor its status.