Research Engineer, Privacy and Anonymization
Posted on Sep 27, 2026
About the Role
You'll build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training at Hud. As a Research Engineer on this team, you'll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You'll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training frontier AI agents.
What You'll Do
- Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case.
- Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
- Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
- Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
- Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields.
- Work with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
What We're Looking For
- Strong proficiency in Python and experience building reliable production data or ML systems.
- Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content.
- Strong experimental instincts and ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
- Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data—and when each is appropriate.
- High attention to detail and ability to reason about subtle leakage paths, edge cases, and adversarial failure modes.
- End-to-end experience building data processing pipelines without a fully prescribed roadmap.
- Hands-on experience with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption.
- Experience working with sensitive data in healthcare, finance, security, or related domains.
- Experience building low-latency or high-throughput ML inference and data-processing systems.