AI Release Quality Gate
An automatic LLM evaluation that runs on every prompt or model change and blocks the release if quality drops.
AI Platform Product Manager · Madrid · Open to EU, UK, Canada, US
9+ years building enterprise data and cloud platforms at IBM, Infosys and Genpact. Most recently I led a GenAI product on Amazon Bedrock from requirements to production and owned 70+ data and AI initiatives. MBA, IE Business School.

faster product delivery with a GenAI product I led on Amazon Bedrock
enterprise data and AI initiatives led across 10 cross-functional teams
annual revenue from the AI and data portfolio I owned
Four hands-on builds for the same European online shop. Two for the AI platform other teams build on, two for AI features customers use. Each one answers: is it good enough to launch, at what cost, and with what risk?
An automatic LLM evaluation that runs on every prompt or model change and blocks the release if quality drops.
An AI agent that resolves order and refund requests, asks a human before risky actions, and is protected against data leaks and prompt attacks.
Routing each request to the cheapest model that is good enough, with a spend-by-team dashboard and an annual savings case.
Answers shopper questions from the product catalogue and reviews with sources, says “I don't know” when unsure, and is tested for conversion.
I map where people lose time before choosing any AI.
Every AI output is evaluated before anyone relies on it.
Success is teams and customers using it every week.