Robot dataset acceptance audit: a verdict before you pay by Jigon YooRobot dataset acceptance audit: a verdict before you pay by Jigon Yoo
Robot dataset acceptance audit: a verdict before you payJigon Yoo
Cover image for Robot dataset acceptance audit: a verdict before you pay
You are about to pay for a robot demonstration dataset. Before you do, you send me the manifest and a sample - or the whole thing - and you get back a verdict on it.
Twelve things get measured, and each one is a pass line you can argue with rather than a score you have to trust: episode completeness; frame count against the declared index; index continuity; timestamp monotonicity; timestamp float drift; joint-state sampling rate against the rate you actually need; camera frame rate against the same requirement; observation and action pairing; joint range conformance against your deployment robot; gripper range conformance against the same; near-duplicate rate across episodes; and required metadata, licence and provenance.
Alongside those you get the sampling disclosure: how many episodes were measured in full versus sampled, stated as a number rather than folded into an average.
Nothing is eyeballed and nothing is scored on vibes. Every failed episode carries the reason it failed. Anything I could not measure is reported as unmeasured rather than folded into a pass. You get the pass line stated up front, the failures listed individually, and a command that reproduces the whole audit on your machine.
This matters because quality filtering typically discards 20 to 30% of a collected dataset. Finding that out after the invoice is expensive; finding it out from a sample is not.
Scope ladder (same as my other channels): $250 - one dataset, up to 500 episodes, manifest and sample. $400 - up to 3,000 episodes, or several datasets audited against one pass line. $650 - the above plus a written acceptance spec you can hand to your vendor as an RFP annex, so the next delivery arrives already measurable.
If your dataset is bigger or messier than the tier you picked, I say so before starting.
The checks run on public, unit-tested tooling - robotdata-pipeline, dataset-curation-dedup and ros2-bag-data-audit at github.com/jigonyoo. Sample projects there use synthetic data and say so.
Starting at$250
Duration1 week
Tags
AI Engineer
Data Engineer
Service provided by
Jigon Yoo Anyang-si, South Korea
Robot dataset acceptance audit: a verdict before you payJigon Yoo
Starting at$250
Duration1 week
Tags
AI Engineer
Data Engineer
Cover image for Robot dataset acceptance audit: a verdict before you pay
You are about to pay for a robot demonstration dataset. Before you do, you send me the manifest and a sample - or the whole thing - and you get back a verdict on it.
Twelve things get measured, and each one is a pass line you can argue with rather than a score you have to trust: episode completeness; frame count against the declared index; index continuity; timestamp monotonicity; timestamp float drift; joint-state sampling rate against the rate you actually need; camera frame rate against the same requirement; observation and action pairing; joint range conformance against your deployment robot; gripper range conformance against the same; near-duplicate rate across episodes; and required metadata, licence and provenance.
Alongside those you get the sampling disclosure: how many episodes were measured in full versus sampled, stated as a number rather than folded into an average.
Nothing is eyeballed and nothing is scored on vibes. Every failed episode carries the reason it failed. Anything I could not measure is reported as unmeasured rather than folded into a pass. You get the pass line stated up front, the failures listed individually, and a command that reproduces the whole audit on your machine.
This matters because quality filtering typically discards 20 to 30% of a collected dataset. Finding that out after the invoice is expensive; finding it out from a sample is not.
Scope ladder (same as my other channels): $250 - one dataset, up to 500 episodes, manifest and sample. $400 - up to 3,000 episodes, or several datasets audited against one pass line. $650 - the above plus a written acceptance spec you can hand to your vendor as an RFP annex, so the next delivery arrives already measurable.
If your dataset is bigger or messier than the tier you picked, I say so before starting.
The checks run on public, unit-tested tooling - robotdata-pipeline, dataset-curation-dedup and ros2-bag-data-audit at github.com/jigonyoo. Sample projects there use synthetic data and say so.
$250