
Manager, AI Benchmarking and Evaluation Research
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in Romania.
• Lead, mentor, and develop a team of researchers and engineers dedicated to assessing AI models utilized in cybersecurity applications.
• Establish the strategy, roadmap, and success metrics for evaluating AI/LLM and agentic systems across various security use cases.
• Design and standardize evaluation methodologies, benchmark datasets, and reproducible testing pipelines based on real SOC workflows.
• Evaluate models for accuracy, robustness, reliability, and operational effectiveness in scenarios involving security analysts.
• Create quantitative and qualitative metrics for AI models and agentic systems in relation to threat detection and response challenges.
• Collaborate with engineering, product, and threat-research teams to convert evaluation findings into scalable enhancements.
• Convey results, trade-offs, and recommendations to both technical and executive stakeholders.
• Hands-on experience in a SOC, with a comprehensive understanding of daily security operations.
• Minimum of 3 years in a management or leadership role, demonstrating a successful history of mentoring and developing technical teams.
• In-depth knowledge of incident response and threat hunting, covering detection, investigation, and remediation processes.
• Broad awareness of the cybersecurity landscape, including attack vectors, defense strategies, and analyst workflows.
• Practical experience with AI, including familiarity with AI/LLM capabilities and insights into agentic systems.
• Ability to define, design, and standardize evaluation methodologies and reproducible testing pipelines.
• Outstanding communication skills, capable of clearly presenting complex technical findings to both technical and non-technical audiences.
• Proven history of delivering results, backed by shareable projects or measurable outcomes.
• Established experience in leveraging AI technologies to enhance decision-making, optimize workflows and processes, boost efficiency, and drive business results.
• Relevant security certifications are advantageous, such as GCIA, GCIH, GCFA, OSCP, or equivalents.
• Experience in building or collaborating with LLM/agentic evaluation frameworks and benchmarking pipelines.
• Programming proficiency in Python and familiarity with data science tools.
• Familiarity with SOC platforms, SIEM/SOAR tools, and detection engineering.
• Exposure to adversarial testing, red-teaming, or robustness assessment of AI systems.
• Understanding of MITRE ATT&CK and its relevance to detection and response workflows.
• Competitive salary
• Stock options
• Private healthcare insurance
• Life insurance
• Training budget
• Opportunity to work with cutting-edge technologies
• Flexible time off
• Team-building activities
• Market-leading compensation and equity awards
• Comprehensive physical and mental wellness programs
• Generous vacation and holiday policies for relaxation
• Paid parental and adoption leave
• Professional development opportunities available for all employees, regardless of position or role
• Employee networks, neighborhood groups, and volunteer opportunities to foster connections
• Dynamic office culture with top-tier amenities
• Great Place to Work Certified™ globally
anni.care
InnoData
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.