Service 04
Generative AI Evaluation & Fine-Tuning Support
Human evaluators rate, rank and review model outputs, RLHF style feedback that supports fine tuning and safety testing.
AIDataTrix runs RLHF and model evaluation services: our own in-house evaluators rate, rank and review your model outputs against your criteria.
What we do
The work itself.
- The evaluators are our own in-house team.
- Ratings come back with Excel and CSV QA trackers alongside the CSV or JSON.
- Rate and rank model outputs
- Review responses against your criteria
- RLHF style structured feedback
- Supports fine tuning cycles and safety testing
To quote
What we need from you.
These are the questions we ask before we can put a number on it.
- What are the outputs: chat responses, summaries, generated images, code, or something else?
- How many items, and how many evaluators should see each one?
- Rate on a scale, rank against each other, or review with written notes?
- Do you have a rubric or criteria, or should we draft one with you?
- Which languages do the outputs use?
- Output format for the feedback: CSV, JSON, or your own schema?
- A one-off batch, or a recurring cycle?
- When do you need the first batch?
Deliverables
What you receive.
- Structured feedback as CSV or JSON
- Excel and CSV QA trackers
- Secure cloud transfer
Assurance
How the work is checked.
All quality checks are handled in house by our own team.
Questions
Asked and answered.
Depending on the engagement: raw and edited video and audio, annotated data as JSON, COCO or CSV, labelled image sets, Excel and CSV QA trackers, or licensed dataset files. Everything moves over secure cloud transfer.
Not yet. Those are on our roadmap. What is in place today: NDAs with all staff and vendors, restricted access to raw data, and secure storage.
We are based in Hyderabad and we work with teams in India and internationally.