Loading jobs…
Loading jobs…
Seekr — Austin Texas
About the Opportunity SeekrFlow is an end-to-end AI development platform used to build, fine-tune, evaluate, and deploy AI systems in production environments. As AI adoption accelerates across enterprises and government agencies, the ability to rigorously evaluate AI models and applications, and to trust those evaluations, has become a critical capability. We’re hiring a Principal Product Manager to own SeekrFlow’s evaluations product area end to end.
This is a senior individual contributor role with broad scope. You will define strategy and drive execution across all evaluation surfaces inside SeekrFlow: base models, fine-tuned models, agents, distilled applications, and beyond. You’ll work in close partnership with other product leaders, ensuring that evaluation capabilities inside SeekrFlow connect seamlessly with scoring, certification, and governance workflows across all Seekr products.
You’ll operate with a high degree of autonomy, partnering directly with Engineering, Research, Design, and GTM to make AI evaluation rigorous, scalable, and customer ready. What You’ll Own Evaluations product strategy & roadmap: Define and own the multi-year strategy and roadmap for SeekrFlow’s evaluation capabilities, covering the full spectrum of model types (base models, fine-tuned models, agents, and distilled applications) and ensuring alignment with platform and business priorities. End-to-end product ownership: Drive evaluation features from discovery and specification through delivery, launch, and iteration, working in close partnership with Engineering and Design.
Own prioritization decisions and hold the line on quality and scope.
Requirements
for each model type and application surface, translating those
into coherent product direction across SeekrFlow. SeekrGuard alignment: Partner closely with other leaders in the technology organization to ensure SeekrFlow evaluation outputs connect directly to SeekrGuard’s risk scoring and evaluation workflows, creating a seamless end-to-end trust pipeline. Research & capability translation: Engage with Research and Engineering teams to identify emerging evaluation techniques and benchmarks, and determine how to turn experimental work into scalable, customer-facing product capabilities.
, workflow gaps, and trust needs. Use those insights to validate direction and sharpen the roadmap. Technical specification: Write clear, detailed product specifications that define evaluation workflows, metrics, API/SDK surfaces, UI behavior, and expected system behavior across SaaS and self-hosted deployments.
Metrics & outcomes ownership: Define success metrics for evaluation capabilities and track adoption, performance, and customer outcomes to drive continuous improvement.
evolved through iteration and technical discovery Proven ability to operate autonomously as a senior IC: setting direction, writing specifications, and driving cross-functional execution Comfortable engaging deeply with engineers, designers, and researchers on technical tradeoffs, system design decisions, and emerging research Experience building for enterprise or government customers in B2B software environments, including multi-stakeholder
and compliance-sensitive contexts Strong product writing skills: able to produce structured, precise specifications
Required Qualifications
8–12+ years of product management experience, with demonstrated ownership of complex, technically deep platform or infrastructure products Strong understanding of AI/ML model evaluation concepts, including benchmarking, bias and reliability testing, task-specific evaluation frameworks, and application types (base models, fine-tuned models, agents) Experience productizing research-driven or experimental technology areas where