Skip to content
WhatsApp

Generate · Annotate · Scale

We shoot and label training data.

Video, audio and image datasets shot to your spec. Video, image, audio and text labelled to your schema. Back as JSON, COCO or CSV.

See what comes back

01Generate

02Annotate

01

Data Collection & Generation

We plan, cast and shoot original video, audio and image datasets to your exact spec, produced end to end in house.

  • Plan the shoot against your specification
  • Cast and direct the people in it
  • Shoot video, audio and image data in house
  • Deliver the material with signed consent on file

Read the service page

Drawing: a shot plan seen from the side, with a camera on sticks, two lights and a subject on a set.
Drawing: an audio waveform with three labelled spans and transcript rows beneath them.

02

Data Annotation & Labeling

Human in the loop labeling of video, image, audio and text, quality checked by our own team.

  • Bounding boxes and segmentation on video and images
  • Transcription on audio
  • Tagging on text
  • Quality checked in house before delivery

Read the service page

Drawing: six dataset cards on a shelf, one of them carrying a licence tag.

03

Off-the-Shelf Data Catalog & Licensing

Pre built, licensable datasets across common and industry specific scenarios, ready to use without commissioning a fresh shoot.

  • Licence a dataset that already exists rather than shooting a new one
  • Common and industry specific scenarios
  • Signed participant consent on everything we recorded
  • Tell us the scenario and volume and we confirm what is available

Read the service page

Drawing: two model responses ranked one above the other, with a rating scale below them.

04

Generative AI Evaluation & Fine-Tuning Support

Human evaluators rate, rank and review model outputs, RLHF style feedback that supports fine tuning and safety testing.

  • Rate and rank model outputs
  • Review responses against your criteria
  • RLHF style structured feedback
  • Supports fine tuning cycles and safety testing

Read the service page

Production

We shoot it, or you send it.

Drawing: a clapper slate, open, with three lines marked on it.

We shoot it

For staged and scenario based datasets we record the video and audio ourselves. Casting, sets, direction, end to end.

Drawing: three file cards moving along a dashed path into a frame with a labelled box in it.

You send it

For annotation only work you send your own video, audio or image files and our team labels them to your spec.

All quality checks are handled in house by our own team.

Data handling

What is in place, and what is not.

We would rather you read this here than ask for it later.

More about how we work

In place today

  • Signed consent from everyone we recordStandard practice on every shoot.
  • NDAs with all staff and vendors
  • Restricted access to raw data
  • Secure storage

Not held yet

  • ISO 27001On our roadmap.
  • SOC 2On our roadmap.
  • GDPR, HIPAA and DPDP alignmentOn our roadmap.

Deliverables

What comes back

  • Raw and edited video
  • Raw and edited audio
  • Annotated data as JSON, COCO or CSV
  • Labelled image sets
  • Excel and CSV QA trackers
  • Structured feedback as CSV or JSON
  • Licensed dataset files
  • Secure cloud transfer
Drawing: a video frame with a person under a segmentation mask and a keypoint skeleton, a dashed box around them, and a second labelled box around a vehicle.
{
  "images": [
    { "id": 1, "file_name": "frame_00412.jpg", "width": 640, "height": 400 }
  ],
  "annotations": [
    {
      "id": 1, "image_id": 1, "category_id": 1,
      "bbox": [198, 84, 172, 272],
      "num_keypoints": 5,
      "iscrowd": 0
    },
    {
      "id": 2, "image_id": 1, "category_id": 2,
      "bbox": [432, 196, 132, 112],
      "iscrowd": 0
    }
  ],
  "categories": [
    { "id": 1, "name": "person" },
    { "id": 2, "name": "vehicle" }
  ]
}
Format example. It describes the drawing above, not a dataset.

Questions

Asked and answered

  • Both. For staged and scenario based datasets we plan, cast and shoot in house. For annotation work you send your own files and we label them to your spec.

  • Depending on the engagement: raw and edited video and audio, annotated data as JSON, COCO or CSV, labelled image sets, Excel and CSV QA trackers, or licensed dataset files. Everything moves over secure cloud transfer.

  • Yes. Signed consent is obtained from everyone we record, as standard practice.

  • Not yet. Those are on our roadmap. What is in place today: NDAs with all staff and vendors, restricted access to raw data, and secure storage.

  • Yes, that is what the catalog service is for. Our catalog is not published online, so tell us the scenario, modality and volume you need and we will confirm what is available and what would need to be shot.

  • We are based in Hyderabad and we work with teams in India and internationally.

Powering Intelligent Data

We shoot it, or you send it.