Senior ML / Evaluation Engineer
Location: Remote from Spain (an indefinite Spanish employment contract)
We are looking for a Senior ML / Evaluation Engineer to help define and implement quality standards for enterprise-grade AI agents and LLM-powered applications. In this role, you will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle. You will work closely with AI Platform Engineers, ML Engineers, and DevOps teams to ensure reliable, measurable, and production-ready AI systems through scalable evaluation, observability, and governance practices.
Project Overview:
Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
Requirements:
Presente su candidatura después de leer los siguientes requisitos de habilidades y cualificaciones para este puesto.
- 5+ years ML engineering or AI platform engineering
- LLM evaluation framework design and implementation
- Custom evaluator implementation for deterministic quality checks
- CI/CD deployment gate design for ML model or agent quality
- AWS AgentCore Evaluation (on-demand mode for CI/CD gates, online mode for production sampling)
- LLM-as-judge evaluator design (built-in AgentCore evaluators — helpfulness, correctness)
- Custom code-based Lambda evaluators (Python — deterministic checks)
- Evaluation levels (TRACE for per-response, TOOL_CALL for per-invocation, SESSION for workflow)
- OTel spans from AWS AgentCore Observability as evaluation input
- Enterprise evaluation standard authoring (mandatory dimensions, pass/fail criteria)
Nice-to-have
- AWS AgentCore Evaluation API hands-on (CreateEvaluation, GetEvaluationResult)
- AWS Bedrock Guardrails for PII detection evaluator integration
- CloudWatch metrics output from AgentCore Evaluation for online mode
Responsibilities:
By continuing you agree to our Terms & Privacy Policy.