MogiMogiJobsPowered by MobiusEngineLet Mogi apply

24-MAG

Remote | Senior Software Engineer - AI Evaluation / Coding Agents — $100–$150/hour

New York, NY · 5 days ago

Contract

$100 to $150 a hour

About the job

We are sharing a specialised freelance opportunity for experienced software engineers to evaluate and improve advanced AI coding systems through rigorous code review, repository-based testing, rubric development, and engineering-quality assessment. Selected professionals will work with coding agents across substantial real-world codebases, review generated implementations, identify technical failure modes, and translate expert engineering judgement into structured evaluation signals. The focus is not primarily on building production applications, but on determining whether AI-generated software is correct, robust, maintainable, and aligned with professional engineering standards. Key Responsibilities AI-Generated Code Evaluation Review code produced by AI coding agents Assess implementations for correctness, robustness, and maintainability Determine whether agents selected appropriate technical approaches Identify subtle implementation errors and engineering weaknesses Explain clearly why generated solutions succeed or fail Rubric & Preference Evaluation Design and refine technical evaluation rubrics Define criteria for assessing coding-agent performance Review and label preference and evaluation data Apply consistent quality standards across repeated assessments Capture nuanced differences between competing implementations Repository & Engineering Analysis Work within substantial real-world software repositories Analyse code changes in their broader architectural context Evaluate repository-level behaviour rather than isolated snippets Review implementation quality using senior engineering judgement Identify recurring failure patterns across coding tasks Evaluation Infrastructure & Workflows Build and maintain pipelines supporting data generation and evaluation Improve infrastructure for collecting and reviewing model outputs Support scalable evaluation and feedback workflows Contribute to reliable processes for repeated technical assessment Help translate qualitative engineering judgement into structured systems Research & Technical Collaboration Collaborate with research and engineering teams Summarise evaluation findings in clear written reports Provide actionable recommendations based on observed model behaviour Help improve evaluation methodologies and coding-agent benchmarks Communicate complex technical issues in concise, structured language Ideal Profile 5+ years of hands-on software engineering experience Strong proficiency in Python, TypeScript, JavaScript, Go, Java, or another major production language Experience working in substantial real-world codebases Strong code-review and technical-analysis skills Ability to assess implementation correctness and maintainability Strong understanding of modern software-engineering practices Experience with GitHub-based development workflows Familiarity with CI/CD and production engineering processes Comfortable analysing unfamiliar repositories and code changes Ability to identify subtle technical issues others may overlook Strong written communication and structured technical reasoning Experience using modern LLMs or AI-assisted coding tools Familiarity with open-source development is valuable LLM evaluation or coding-agent experience is advantageous Experience with RLHF, preference data, rubric design, or post-training is useful but not required Engagement Details Independent contractor engagement Remote — North America only Compensation: $100–$150/hour Candidates should provide a specific hourly rate expectation Flexible commitment of approximately 20–40 hours per week At least 6 hours of Pacific Time overlap per day is required Expected project duration is approximately 3 months Start is as soon as possible Evaluation process includes an approximately 25-minute AI interview A practical code and AI-evaluation exercise of approximately 30 minutes follows Final stage includes an approximately 20-minute hiring manager interview The practical exercise focuses on reviewing AI-generated code rather than competitive programming or algorithm puzzles Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, repositories, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy