RemoteJobs.org mascotRemoteJobs.org
Remote JobsCompaniesAPISign inPost a Job
RemoteJobs.org mascotRemoteJobs.org

Find your dream remote job. Browse thousands of remote positions from top companies worldwide.

Job Categories

  • General
  • Programming
  • Design
  • Marketing
  • Sales
  • Customer Support

Resources

  • Browse Jobs
  • Companies
  • Post a Job
  • For Developers
  • Blog

Company

  • About Us
  • Contact
  • Privacy Policy
  • Terms of Service
© 2026 RemoteJobs.org. All rights reserved.
    ← Back to all jobs
    Mindrift

    Freelance Agent Evaluation Engineer

    Mindrift
    Contract
    Verified Remote
    RemoteUSD 50 - 50ProgrammingToday

    About this role

    Please submit your CV in English and indicate your level of English proficiency.

    Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

    We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

    You'll create challenging tasks and evaluation criteria within realistic simulated environments:

    • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history

    • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent

    • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient

    • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

    What this is NOT:

    • Not data labeling

    • Not prompt engineering

    • Not writing code from scratch - the agent writes most of the code; you guide and evaluate

    What we look for:

    • 5+ years in software development

    • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis

    • Experience writing tests (functional, integration)

    • English proficiency - B2+

    Why this is hard:

    Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

    Requirements and benefits

    Educational qualifications

    • A Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems or other related fields.

    • Bachelor’s degree is accepted if only candidate has 5 years of experience in the field.

    Academic and/or Professional Experience

    Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles

    How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

    Compensation:

    Paid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks — up to the equivalent of $50/hr. Because payment is per task, a faster pace raises your effective hourly rate.

    Why this freelance opportunity might be a great fit for you?

    • Take part in a part-time, remote, freelance project that fits around your primary professional or academic commitments.

    • Work on advanced AI projects and gain valuable experience that enhances your portfolio. - Influence how future AI models understand and communicate in your field of expertise.

    About Mindrift

    Mindrift
    Mindrift

    Mindrift, powered by Toloka, is a platform that connects domain specialists and experts with AI project opportunities from major tech innovators. The company focuses on post-training and evaluation of frontier AI models by creating domain-specific reinforcement learning environments, tasks, and evaluation frameworks. Mindrift enables experts across fields—including mobile development, management consulting, physics research, and other domains—to shape how next-generation generative models learn and perform by converting real-world expertise into structured learning environments. The platform operates on a project-based model, allowing contributors to work on diverse AI initiatives including mobile app development, consulting domain training, physics problem design, and other specialized tasks for leading technology companies.

    Hiring remote talent?

    Reach active remote job seekers from $149.

    Related Jobs

    Senior Software Developer

    Parallels · CAD 120,000 - 145,000

    Senior SAP Solution Architect – SAP TM Transformation - Remote with some travel

    Simple Software Solutions Group

    FULL-STACK DEVELOPER

    Human Power BG