Loading jobs…
Loading jobs…
Foursquare Labs, Inc. — Washington D.C
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.
We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers. About the Program: Innodata's Federal Practice builds the trusted data layer for critical infrastructure Trust Safety work.
Partnering with a leading systems integrator, we're delivering a modern, governed data services platform in a secure federal (IL4) environment. Over an intensive 20-week phase, you'll help stand up a data services storefront, a DataCard governance framework, synthetic data integration, and Databricks write-back capabilities.
About The Role
: As the QA/Evaluation Lead, you'll own quality and evaluation across the platform. You'll design the evaluation framework that measures whether our data services and outputs meet the bar, build repeatable test and validation processes, and give the team an objective read on readiness at each milestone. Partnering with the Delivery Owner and engineering leads, you'll turn quality from an afterthought into a measurable, demonstrable strength.
It's a role for someone who thinks rigorously about evaluation and takes pride in evidence-backed quality.
Key Responsibilities
: Design and own the inter-annotator agreement (IAA) methodology for the Phase 1 demonstration corpus — metric selection (Cohen's kappa, Fleiss, Krippendorff's alpha), sampling design, adjudication workflow, and agreement thresholds Define evaluation framework architecture: test and evaluation plans, IAA targets, drift detection gates, and model performance metrics per SOW Section 2.9 Configure and operate sampling-based quality control across the self-service and white-glove annotation paths during Phase D corpus production Design and implement confidence-threshold escalation routing from automated annotation to senior-annotator adjudication Validate quality scoring and IAA computation within the Innodata data layer Support AI Solutions Engineer on evaluation design for SAM 2 and Frontier model API validation — define what 'good enough' looks like quantitatively Produce evaluation framework documentation for the Phase 1 NPP closeout package, including per-DataCard documentation with the SA Must-Have
Qualifications
: Bachelor's degree in Statistics, Data Science, Computer Science, or related quantitative field required; Master's degree preferred. Equivalent experience may substitute for degree on a 2-for-1 basis.
: Prior DoD or IC data quality program experience CVAT or equivalent annotation platform QC workflow configuration Drift detection and model monitoring methodology Experience with FMV / video annotation quality standards The expected hourly salary range for this position is $45 to $50 p/hour, based on experience, skills, and
. Note to Candidates: This role is not a project manager with QC
Responsibilities
— it is a methodology expert who owns the intellectual framework behind data quality on a federal AI program.