QA Evaluation Engineer
Location: Remote (must reside in Ukraine or Poland)
About Napster:
Founded on the principle of democratizing access—first to music in 1999, now to creative expertise in 2026—Napster has consistently been at the forefront of transformational technology shifts that expand markets and empower users. The company’s latest platform turns passive consumers into active creators, providing the interface layer between foundation AI models and human creativity. For more information, visit napster.com.
Napster is looking for a QA Evaluation Engineer who thinks in rubrics and calibration, and treats “did the AI actually do its job well?” as a question with a measurable answer. This is not a check-the-box QA role — you’ll own how we prove our conversational agents are grounded, on-persona, and correct at a scale no manual pass can reach. You’ll build the evaluation capability from the ground up, make judge-and-rubric decisions that outlast any single release, and turn fuzzy product judgment into scores the whole engineering org can trust. As our agents grow more capable — multi-step and tool-using — this role grows with them, into evaluating agentic behavior itself.
What You’ll Own
Evaluation Framework & Harness
- Own an LLM evaluation framework (like Promptfoo) end to end — judge configuration, execution model, and its evolution as the product changes
- Define how evaluation runs across layers: deterministic checks for what’s checkable (schema, tool calls, latency, language) and model-graded judging for what isn’t (groundedness, task success, tone/persona)
- Wire evaluation results into the existing reporting and defect-tracking workflow so failures surface where engineering already works
- Validate the UI surfaces our agents are delivered through — conversation experience, configuration, and knowledge-source management
Golden Datasets & Rubrics
- Curate labeled golden datasets for a representative cohort of agents, including negative cases — when the answer isn’t in-source, the agent should decline, not fabricate
- Author and maintain the scoring rubrics that make automated evaluation trustworthy, in partnership with PMs who supply product-side judgment on what “good” means per agent
- Standardize a per-agent onboarding scaffold — fixture corpus, acceptance-criteria rubric, outcome checks — so evaluation effort doesn’t scale linearly with the catalog
Judge Calibration & Trust
- Treat LLM-as-judge validation as first-class: verify the automated grader agrees with human labels on held-out sets, and keep that calibration current as prompts and knowledge sources change
- Reason explicitly about inter-rater and judge–human agreement — the scores are only as trustworthy as this check
- Guard against the failure modes that quietly break automated grading: judge bias, tautological rubrics, and drift
Coverage Strategy at Scale
- Design a tiered coverage strategy rather than exhaustive per-agent testing: deep evaluation of a representative cohort, shared checks that apply fleet-wide, and sampling to detect drift across the long tail
- Run outcome-focused UAT at release milestones — where the judgments you make become the golden-set content the framework runs on
- Build rubrics and harness components generic enough to carry across product lines, so scores stay comparable and the capability travels to the next initiative
What We’re Looking For
Required
- 3+ years in QA or quality engineering, plus hands-on experience evaluating LLM or conversational-AI outputs — groundedness, task success, tone/persona, structured-output correctness. That evaluation experience can be recent; it’s a young field, and we weight depth over years here
- Direct experience with an LLM evaluation / model-graded framework (e.g. Promptfoo) and the LLM-as-judge techniques it enables
- Experience curating labeled / golden datasets and designing scoring rubrics, with an eye for inter-rater and judge–human agreement
- A strong general QA foundation: test planning and strategy, test design from requirements and acceptance criteria, and deliberate thinking about coverage and risk
- Breadth across testing types — functional, regression, exploratory, and UI / application testing, including the judgment to make release-gating calls
- Sharp defect practice: reproduction, isolation, clear and actionable bug reports, and ownership of the defect lifecycle
- A probabilistic view of evaluation — comfortable in pass rates and tolerance bands rather than strict pass/fail
- Comfort partnering with Product to turn fuzzy judgment about “good” into explicit, measurable criteria
Preferred
- Familiarity with RAG internals — chunking, embeddings, retrieval — and where grounding typically breaks
- Working understanding of LLM-as-judge pitfalls and calibration
- Multi-turn / conversational evaluation experience (simulated users, transcript grading)
- TypeScript and/or Python; familiarity with CI and QA reporting / test-management tooling (e.g. Allure, Jira, TestRail) or equivalents
- Exposure to real-time / voice channels (WebRTC, SIP/VoIP) — relevant as evaluation extends across transports
- A track record of standing up an evaluation capability from scratch and scaling it
Our Culture
- Impact: Play a crucial role in our growth journey.
- Culture: Join a vibrant team valuing creativity and collaboration.
- Growth: Thrive in a fast-paced, dynamic environment.
- Reward: Enjoy competitive compensation, equity opportunities, and comprehensive benefits.
- Ready to shape our future? Apply now and be part of something extraordinary!
We’re looking for more forward-thinking, collaborative people to be part of our innovation journey and mission to push the boundaries of technology. If you’re ready to help us achieve this vision, we’d love to hear from you! At Napster Corp, we're looking for people who are invigorated by our values and driven to change the world, not those who simply check off boxes.
Napster Corp embraces a diversity of backgrounds and experiences and provides equal opportunity for all applicants and employees. We strive to build a company that reflects a global audience.