... - English (India) Type: Contract Compensation: $12/hour Location: Remote Duration: Up to 6 months Commitment: 10+ hours/week Role Responsibilities - Evaluate AI model output lyrics, voice generation, and other standards in music. - Score the quality of musical training data to enhance model performance. - Collaborate with ...
... into actionable model or prompt improvements.- Establish AI governance documentation, audit trails, and safety reports.- Embed evaluation and safety checks into AI release pipelines with engineering teams.- Track relevant AI governance and regulatory frameworks, including NIST AI RMF and India/EU AI https://jobeax.com/link/XZWIbwdqxtvXofHu ...
... anymore — they're asking why some conversations succeed while others fail. Understanding that difference is one of the most important challenges in enterprise Voice AI today. Blue Machines runs production systems handling millions of conversations for India's largest airlines, banks, mutual funds, and retail chains. Not prototypes. ...
... Build grading criteria that define correct answers for scientific tasks. - Calibrate tasks against frontier models, ensuring tasks ship only when strong models fail more often than succeed. - Work independently and asynchronously to meet deadlines while improving AI model performance . Qualifications Must-Have - PhD in mathematics ...
... Build grading criteria that define correct answers for scientific tasks. - Calibrate tasks against frontier models, ensuring tasks ship only when strong models fail more often than succeed. - Work independently and asynchronously to meet deadlines while improving AI model performance . Qualifications Must-Have - PhD in mathematics ...
... US-based, actively practicing physicians (MD/DO) to evaluate clinical EHR vignettes for a medical AI evaluation program. General and internal medicine are our main focus. While each project involves unique tasks, contributors may: - Evaluate clinical EHR vignettes paired with a question, a proposed answer, and a distractor ...
... development tools (Claude Code, Cursor, Copilot or similar) as part of your normal workflow for implementation, test generation, and coverage expansion. - Build evaluation harnesses for AI and agent-based features, where outputs are probabilistic and traditional assert-equals testing does not apply. - Apply AI to anomaly, regression, ...
We are looking for advanced Physical Sciences experts to support a scientific computing and AI evaluation project involving realistic, terminal-based scientific tasks. The project focuses on translating real-world scientific workflows into reproducible computational tasks and evaluating whether AI systems can correctly ...
... clinical image interpretation skills to the evaluation of emerging AI diagnostic technology. Your feedback can help assess the performance and clinical utility of AI-generated dermatological diagnoses against professional standards. Apply to participate and bring your dermatology expertise to the evaluation of AI-powered diagnostic ...