24-MAG
Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Contract
About the job
We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research. Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows. Key Responsibilities Code Curation & Solution Development Curate high-quality code examples for model training and benchmarking Develop precise solutions to software-engineering tasks Correct and...