Skip to content
RCCData Services
RCCData Services

Data infrastructure for AI agents

Task-oriented datasets for coding agents, interactive environments, frontier reasoning, safety, and multilingual evaluation.

Explore Datasets
datasets in the catalogue
10
built from long-horizon tasks
6
Samples on request
  1. 01tasks/fix-failing-build/task.toml

    schema_version = "1.4"[metadata]category = "software-engineering"[agent]timeout_sec = 900.0[environment]network_mode = "no-network"
  2. 02Task files

    instruction.mdenvironment/tests/solution/
  3. 03harbor run

    1. 01
    2. 02
    3. 03
  4. 04Verifier result

    verifier/reward.txt
  • RL Environments
  • Terminal & Coding Tasks
  • Long-Horizon Tasks
  • Interactive Agents
  • Frontier Reasoning
  • Safety & Alignment
  • Multilingual Evaluation

01Dataset catalogue

Task-oriented datasets for training, evaluating and benchmarking AI systems, from single terminal tasks to work that runs for hours.

10 of 10 datasets

02Solutions

Start from the evaluation problem, then narrow to the data. Each area links to the catalogue entries that serve it.

A

From a single command to a whole repository.

Coding & Software Engineering

For teams evaluating terminal-based workflows, repository-level tasks, code transformation, and software engineering agents.

  • Coding agent evaluation
5 datasets
B

Hours, not turns.

Long-Horizon Agent Evaluation

For teams measuring whether agents sustain correct progress across extended work, from two-to-three-hour engineering marathons to frontier tasks that take an agent around twenty hours.

  • Long-horizon evaluation
6 datasets
C

Agents that act, observed step by step.

Interactive Environment Evaluation

For teams working with agents that interact with applications, environments, tools, or changing task states.

  • RL environment evaluation
  • Computer-use agents
1 dataset
D

The upper end of agent capability.

Frontier Reasoning & Knowledge

For teams evaluating advanced problem-solving and knowledge-intensive agent behavior.

  • Frontier science & knowledge
3 datasets
E

Behavior under constraints, in every supported language.

Safety & Multilingual Evaluation

For teams studying agent safety, instruction alignment, and performance across supported languages.

  • Agent safety & alignment
  • Multilingual evaluation
2 datasets

03Methodology

A reference structure for task-oriented evaluation data, readable by researchers and procurement teams alike. Use it to judge whether any dataset gives you what you need to measure an agent.

Stage 01 of 05

Task definition

Define the objective, starting conditions, constraints, and expected outcome.

instruction.md · task.tomlHarbor · illustrative
instruction.md
what the agent is asked to do
[task] name
"example/fix-failing-build"
[metadata]
category, difficulty, tags
time estimate
expert_time_estimate_min = 60

These stages describe a robust evaluation-data workflow in general terms. How a specific RCC dataset maps to each of them is something to confirm with the team as part of a sample request.

04Why RCC

A dataset is only useful if it measures what you care about. The catalogue and the request process are built around that question.

  1. 01

    Discovery by task domain

    The catalogue is organised by what you need to evaluate: coding, interactive environments, frontier reasoning, safety and multilingual work.

  2. 02

    Samples requested per dataset

    Ask for samples of exactly the datasets you are considering, one at a time or several in a single request.

  3. 03

    Clear about task scope

    Each entry says what kind of task it covers, and separates what is documented from what is confirmed with you directly.

  4. 04

    Technical conversations

    Talk through environment, tooling and evaluation requirements with the team before you commit to anything.

  5. 05

    A direct route to a decision

    From catalogue entry to sample to a judgement about fit, without a long qualification survey in between.

example/fix-failing-build/task.toml
schema_version = "1.4"artifacts = ["/app"] [task]name = "example/fix-failing-build" [metadata]category = "software-engineering"difficulty = "hard"tags = ["terminal", "python"] [agent]timeout_sec = 900.0 [verifier]timeout_sec = 300.0environment_mode = "separate" [environment]network_mode = "no-network"cpus = 2memory_mb = 4096
An illustrative task in the Harbor format: the instruction, its configuration and environment, the verifier that scores it, and a reference solution. A dataset's actual files are part of what to review in a sample.

05About

About RCC Data Services

RCC Data Services helps research and engineering teams working on coding agents, reinforcement learning environments, long-horizon tasks, benchmark evaluation, frontier reasoning, agent safety and multilingual capability source task-oriented data, and assess a sample before they commit.

Contact

The quickest route to the team is a sample request, or a short note about the requirement you're working on.

06Sample requests

Tell us which dataset you're exploring. We'll use your request to understand what you're looking for and follow up about sample availability.

What happens next

  1. 01Your request is recorded with the datasets you selected.
  2. 02The team follows up about sample availability.
  3. 03You assess the sample against what you need to evaluate.

Four fields at most. No phone number, budget or long questionnaire.

We use these details only to respond to this request and follow up about the datasets you selected.