Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Trainer Model Test and Measurement Engineer at DEF CON

Evaluates AI models by building labeled datasets, designing audits, measuring performance, and validating scoring and generative AI outputs.

Senior Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

ABOUT DEFCON AI

RESILIENCE IN THE FACE OF DISRUPTION. DEFCON AI is an insights company that leverages artificial intelligence, mathematical optimization, data analytics, and software engineering for resilient optimization of complex systems.

In today’s dynamically changing world, DEFCON AI’s technology aligns outcomes with operational goals, better decision making, and empowers customers to anticipate assess, and mitigate the impacts of disruptions.

Be the independent voice that keeps the whole program honest — your evaluation is the standard everyone else is held to.

About the Role

You’ll join the analytics and AI engineering team behind a system that genuinely matters: an AI-assisted platform that brings together records from dozens of disparate data sources, resolves them to the correct individual, highlights what analysts should review first, and provides transparent, explainable recommendations that users can trust. Operating within a secure government cloud environment, the platform tackles complex challenges in AI, data integration, and decision support where quality, trust, and accountability are mission-critical.

As a Model Test & Measurement Engineer, you’ll own the evaluation framework that helps ensure those systems perform as intended. You’ll build and maintain labeled ground truth datasets, design statistically sound audit and sampling methodologies, measure model performance across releases, and create the evidence packages that support deployment decisions. You’ll independently validate both scoring and generative AI capabilities, helping the team understand not only whether a model works, but how confidently its outputs can be trusted.

This is a role with genuine influence. Your assessments will inform release decisions, drive improvement efforts, and provide the objective evidence customers rely on when evaluating system performance. You’ll work closely with data scientists, AI engineers, and technical leadership while maintaining the independence needed to provide clear, defensible evaluations. You do not build the models you validate, and you do not approve thresholds or release authorization against your own evidence.

If you’re energized by measurement, validation, and making complex AI systems more trustworthy, this is an opportunity to have an outsized impact on both the technology and the mission it supports.

This is a fully remote role with occasional travel to DEFCON AI headquarters, customer sites, and partner facilities as needed.

Key Responsibilities

  • Construct labeled ground truth for model evaluation

  • Design and run sampled audits

  • Gate releases against model version; maintain version inventory, evaluation records, and rollback triggers

  • Run drift and override review

  • Produce human-oversight and fairness / disparate-effect evidence

  • Maintain independence from the build roles: produce evaluation evidence, but do not build the models, approve thresholds, or authorize releases against that evidence

  • Measure workflow improvement: review time, throughput, backlog movement, and override and rework rates

  • Design evaluation-phase QC: sample selection that does not mix populations, the unit of review, and fair comparison when methods or searched sources differ

  • Define the measurement events other teams must emit, and confirm the review workspace emits workflow events from first use

Required Qualifications

  • 5+ years with model validation as a named responsibility, not a side task

  • Direct experience with ground-truth construction and sampling design

  • Strong Python and SQL; comfort working independently from the teams whose models you evaluate

  • US Citizenship Required

  • Active US Secret clearance

  • Elevated personnel security requirements apply to portions of this work and are discussed during screening

Preferred Qualifications

  • Regulated-industry or government model-risk background

  • Experience with fairness / disparate-effect testing and human-oversight documentation

  • NIST AI RMF or comparable practice

  • Active Top Secret clearance

What Success Looks Like

  • Evaluation evidence a customer can rely on, independent of the teams that build the models

  • Releases gated against a clear version and evaluation record

  • Workflow metrics (review time, throughput, backlog, override/rework rates) that give the program an honest read on whether it’s working

What We Offer

  • A fully remote, results-based environment

  • Competitive salary, bonus, and equity package

  • 100% employer paid, comprehensive health insurance including medical, dental, and vision for you and your family

  • Unlimited PTO, with your manager’s approval

  • Flexible work environment where you manage your work day

  • 14 weeks of fully-paid parental leave

Salary Range: $150,000–$190,000. This represents the typical salary range for this position based on experience, skills, and other factors.

We’re an Equal Opportunity Employer: You’ll receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or on the basis of disability.

Applicant Data Disclosure

By submitting an application, you acknowledge that Defcon AI uses third-party service providers to facilitate its recruitment and hiring processes. These providers include applicant tracking systems, candidate verification platforms, and fraud detection tools (collectively, “Hiring Platforms”). Your application materials, including your résumé, cover letter, work samples, responses to application questions, and any other information you submit, may be transmitted to and processed by these Hiring Platforms for the following purposes:

  • Managing and administering your application throughout the hiring process;
  • Verifying the accuracy and authenticity of application materials, including by cross-referencing information you provide against publicly available sources and proprietary databases;
  • Identifying indicators of potentially fraudulent, fabricated, or materially misleading application content, including but not limited to discrepancies between submitted materials and publicly available professional profiles, geographic anomalies, and fabricated work histories.

Applications that are flagged through this process as containing indicators of fraud or material misrepresentation may be declined from further consideration. If you have questions about the status of your application or the evaluation process, please contactrecruiting@defconai.com.

Defcon AI requires its Hiring Platform providers to process your information solely for the purposes described above and in accordance with applicable law. Your information will be retained only for as long as necessary to fulfill these purposes and any applicable legal obligations, after which it will be deleted in accordance with Defcon AI’s data retention policies.

For more information about how your data is used, please refer to our Privacy Policy and Applicant Privacy Notice .

Read the full description
Trainer AI Safety Red Teamer Expert

Tests AI systems adversarially to identify safety risks, vulnerabilities, and harmful model behaviors.

Senior Posted 1 day ago Himalayas
What this role involves
About the jobMercor connects elite creative and technical talent with leading AI research labs.
Read the full description
Trainer Senior Software Engineer, AI Training - UK

Evaluates and provides expert feedback on coding-agent interactions to improve modern AI programming systems.

Senior Remote Posted 3 days ago Himalayas
What this role involves
Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week Check out this Loom video for more details: We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
Read the full description
Trainer Senior Software Engineer, AI Training - Canada

Evaluates and provides expert feedback on coding-agent interactions to assess and improve AI training quality.

Senior Remote Posted 3 days ago Himalayas
What this role involves
Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week Check out this Loom video for more details: We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
Read the full description
Trainer Senior Software Engineer, AI Training - Ireland

Evaluates and provides expert feedback on coding-agent interactions to improve modern AI software development tools.

Senior Remote Posted 3 days ago Himalayas
What this role involves
Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week Check out this Loom video for more details: We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
Read the full description
Trainer Senior AI Trainer

Trains AI models by evaluating and providing feedback on model outputs as a remote senior contractor.

Senior Remote Posted 8 days ago Himalayas
What this role involves
Role Title: Senior AI Trainer Role Type: Contractor, Remote Location: Northern America and Europe.
Read the full description
Trainer Pure Mathematics Specialist – Freelance AI Trainer Project

Provides mathematical expertise to train and improve AI models through problem-solving and feedback on theoretical reasoning tasks.

Senior Remote Posted 17 days ago Himalayas
What this role involves
Are you a theoretical mathematics expert eager to shape the future of AI?
Read the full description
Trainer AI Mentor (Independent Contractor) – Business Leader Programs

Mentors business leaders through AI-focused educational programs, providing guidance and feedback on AI concepts and applications.

Senior Remote Posted 18 days ago Jobicy AI
What this role involves
About Us Udacity is now an Accenture company, and exciting things are happening! 🚀 We are on a mission of forging futures in tech through radical talent transformation in digital...
Read the full description
Trainer Agentic AI Technical Mentor – Independent Contractor (US Canada, Europe, MENA, APAC)

Mentor and guide learners through agentic AI concepts and practical applications as an independent contractor across multiple global regions.

Senior Remote Posted 18 days ago Jobicy AI
What this role involves
About Us Udacity is now an Accenture company, and exciting things are happening! 🚀 We are on a mission of forging futures in tech through radical talent transformation in digital...
Read the full description
Trainer Python Engineer, AI Coding Agent Evaluator

Evaluates and rates AI coding agent outputs to improve model performance and quality.

Senior Remote Posted 23 days ago Himalayas
What this role involves
Senior AI Interaction Evaluator (Codex / Claude Code)Contract | $100–$200/hour | 10–20 hrs/week | Start ASAP (through early May) Check out this Loom video for more details!
Read the full description