Skip to main content
Gramian Consulting Group logo

Scientific Chemistry (Python, AI Evaluation)

Gramian Consulting Group
2 hours ago
Contract
Remote
Worldwide
AI Trainer Jobs – Train AI Systems In Your Area Of Expertise

Scientific Chemistry (Python, AI Evaluation)

Company Gramian Consulting Group
Department Talent Solutions
Location Remote (Bangladesh, Brazil, Colombia, Egypt, Nigeria, India)
Type Contract
Remote Yes
Posted 2026-09-19

About Us

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

Role overview

We are looking for a scientific computing professional to contribute to an advanced AI training initiative focused on developing rigorous STEM coding datasets. The project involves creating and validating complex chemistry problems that help train and evaluate frontier AI models. You will be responsible for scientific problem design, Python implementation, and rigorous quality validation, ensuring that every task meets strict standards for scientific correctness, determinism, and evaluation performance.

CONTRACT: Short-term freelance

COMMITMENT: Hourly - up 40 40h/week

LOCATIONS: Remote - Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Pakistan, Indonesia, Kenya, Nigeria, Turkey, Vietnam

NOTE:Must have published research or academic project experience in a STEM domain.

Responsibilities

  • Design scientific coding tasks consisting of one main chemistry problem and at least 3 logically connected, progressively structured sub-problems.
  • Develop verified golden solutions in Python with complete unit test coverage.
  • Create discriminative test cases that distinguish correct from incorrect AI-generated outputs.
  • Execute quality control validation through the Central Task Platform (CTP), including Tier 1 structural checks and Tier 2 quality rubrics.
  • Iterate on task specifications, implementations, and test cases based on QC feedback.
  • Optimize tasks to meet Pass@K evaluation criteria across multiple LLM judges, including GPT, Gemini, and Nemotron.
  • Ensure scientific correctness, determinism, well-posedness, and adherence to project quality standards.
  • Maintain high first-submission quality and minimize rework, targeting consistent L1 approval.
  • Participate in review meetings, feedback sessions, and project standups during required overlap hours.

Requirements

  • Master's degree or PhD in Chemistry.
  • Strong Python programming skills with experience in scientific computing.
  • Experience using scientific Python libraries such as NumPy, SciPy, SymPy, or relevant domain-specific tools.
  • Ability to formulate rigorous scientific problems with clear constraints and expected outputs.
  • Experience writing or implementing scientific solutions with unit tests and deterministic outputs.
  • Prior experience in AI data annotation, scientific research, or scientific writing.
  • Familiarity with LLM evaluation frameworks, coding benchmarks, or AI-generated code assessment.
  • Published research or academic project experience involving chemistry or another STEM discipline.