AI & ML datasets

Move training data and checkpoints between clouds, clusters and labs as one package.

A training corpus is millions of small files; a checkpoint is a handful of very large ones. Both have to move between wherever the data is and wherever the GPUs are, often across regions or providers. RIPTON CLOUD moves either as a single package.

Plan limits are those published on the pricing page. Benchmark figures come from RIPTON CLOUD's published EC2 test at roughly 150 ms latency and 3% packet loss. Results on your route depend on your network, endpoints and storage. RIPTON CLOUD does not claim any industry certification on this page.

Evaluated transfer path
01

Data lake or lab storage

02

RIPTON Desktop or CLI

03

RIPTON node near the cluster

04

Cluster or object-storage staging

Dataset in, the same tree on the compute side, with a completion record for the run log.

Data gravity

The GPUs are in one place. The data is in another.

Compute is bought where it is available; data lives where it was collected or licensed. Every training run starts by moving one to the other, and the move is often the slowest, least observable step in the pipeline.

01

Millions of small files

Image, audio and text corpora are file-heavy. Per-file negotiation in ordinary tools turns a fast link into hours of overhead.

02

Cross-region and cross-provider routes

Latency between regions and providers is where TCP throughput falls away, and egress charges make retries expensive.

03

Partner labs and vendors

Annotation vendors and research partners send and receive data without access to your cloud accounts or cluster.

Workflows

Where the transfer path matters most.

Image sets · audio · text shards

Corpus to the training cluster

Package the dataset tree once and send it to a node beside the cluster. The structure arrives intact for the data loader.

Checkpoints · eval sets · exports

Checkpoints between sites

Move multi-gigabyte checkpoints and evaluation artefacts between clusters or back to the lab with the CLI or a schedule.

Labelling · red-team sets · audits

Data in and out of annotation vendors

Give a vendor a Shared Inbox for returned labels and send them raw batches as packages, with no account on your infrastructure.

RIPTON CLOUD in the workflow

What an ML team gets

The product earns its place by moving prepared data between endpoints—not by claiming ownership of every system around it.

File-heavy corpora as one package

In the published test, 100,000 small files finished in under a second against nearly three minutes for the TCP baseline at 150 ms and 3% loss.

CLI and API for the pipeline

From Professional, a CLI, REST API, schedules and webhooks so transfers are a step in the run, not a ticket.

Your nodes, your storage

Transfer nodes run on infrastructure you control, cloud or on-premises, beside the storage they read and write.

Encrypted on every packet

AES-256-GCM or ChaCha20-Poly1305 with a forward-secret key per session and replay protection. There is no plaintext mode to misconfigure.

No generic speed promise

Throughput depends on your route, endpoints, and storage. The published benchmark (a 1 GB file 7.5x faster than the TCP baseline at 150 ms and 3% loss) tells you where the difference appears; your own route tells you how much.

Product boundary

A transfer step, not a data platform.

RIPTON CLOUD moves prepared datasets and artefacts between endpoints. It does not version datasets, serve features or orchestrate training.

What it covers

  • Datasets and checkpoints sent as packages
  • CLI, API and scheduled transfers
  • Inbound data from vendors and partners
  • Completion records for the run log

What stays with your stack

  • Dataset versioning or lineage
  • Feature stores or model registries
  • Training orchestration
  • Object-storage replication services

Evaluation path

Test it on one real corpus

Pick the dataset move that delays a run today and measure both tools on the same route.

  1. 01

    Choose the move

    Lake to cluster, cluster to cluster or vendor to lab. One route, one real dataset.

  2. 02

    Move the corpus

    The whole tree, so file-count overhead is measured. Record elapsed time and restarts.

  3. 03

    Move a checkpoint

    One large file the same way, so both workload shapes are covered.

  4. 04

    Decide on numbers

    Elapsed time, retries and any egress incurred by retries, written down.

FAQ

Questions technical buyers ask.

Does it read from S3 or GCS directly?

Transfer nodes read and write the storage they are attached to. Most teams run a node on an instance beside their object storage and stage to it; the exact layout is part of the evaluation.

Can we script it?

Yes. From Professional, a CLI and REST API, plus schedules, watch folders and webhooks.

Is there a file-size or file-count limit?

No file-size limit from Professional up; 10 GB per file on Free and 100 GB on Starter. Packages can contain any number of files within the plan's monthly allowance.

Will it be faster than rclone or aws s3 cp?

On a short, clean link within one region, probably not by much. Across regions and providers, where latency and loss rise, the published benchmark shows where the difference appears. Measure it on your route.

Technical evaluation

Get the data to the GPUs.

Start a free workspace, or bring the dataset move that delays your runs and we will help you measure it.