AI & ML interests

Private evals, RL environments and training data for AI agents on real creative jobs. First up, video editing.

Recent Activity

Organization Card
TensorTest: test and train your agents on real creative jobs.

Test and train your agents on real creative jobs.

Private evals for AI labs, plus RL environments and training data. First up, video editing.

VIDEO EDITING · TIMELINEBENCH

Agents say they’re done. Most aren’t.

93%Ended with the agent claiming success
14.0%Resolved [11.7, 16.4]

TimelineBench gave 16 agent setups 56 real editing jobs. Each is a whole job: raw footage, audio and a brief in, a finished cut out. A run resolves only if it passes hard checks on delivery, content and the brief, and a quality test calibrated on blind judgments by 43 professional editors.

Each square is one agent run. 836 of 896 runs ended with the agent claiming success; 125 (14.0%) resolved.
Each square is one agent run. Source: arXiv:2609.35143.

What we build

Private evals

Available now

We run your unreleased checkpoint or agent on our jobs and send a failure report. Results stay private.

RL environments

Scoped with your team

Our jobs, packaged for your sandbox, with our checks exposed as the reward.

Training and eval data

Commissioned

New jobs on commissioned footage with training rights, plus professional editors’ trajectories, preference pairs and rubrics.

Tell us which creative job to measure next. Terms are scoped per engagement.

The three founders wrote TimelineBench.

Two of us were Research Fellows at Microsoft Research India. Papers at NeurIPS, ACL Findings, EACL, Interspeech and CSCW.

  • Pavan Kalyan Tankala CEO
  • Nirmit Arora CTO
  • Gunin Gupta CBO
  • Y Combinator W26

Want your next checkpoint run privately before release?

TensorTest (formerly Ritivel) · San Francisco

models 0

None public yet