Job Description
We are seeking a Principal Engineer, AI Safety to lead the architecture and technical strategy for safeguarding large language models (LLMs) and multimodal AI systems from misuse. This role focuses on developing scalable defenses against jailbreaks, prompt injection attacks, harmful outputs, and policy violations while ensuring models behave safely and reliably in production environments. The position combines AI safety, machine learning, systems engineering, and evaluation frameworks to improve model robustness and alignment.
Key Responsibilities
Design and lead the development of model-level defenses against jailbreaks, prompt injection attacks, and policy violations.
Develop and implement evaluation frameworks, adversarial testing methodologies, and stress-testing pipelines to identify safety weaknesses before deployment.
Define technical strategy for mitigation approaches, including:
Safety-focused fine-tuning
Prompt shielding and guardrails
Output filtering and post-processing techniques
Collaborate with researchers and adversarial testing teams to convert emerging threats into measurable evaluations and production safeguards.
Build and scale human-in-the-loop review systems to identify toxic, biased, unsafe, or policy-violating model behavior.
Monitor advancements in AI safety research, adversarial attack methods, and jailbreak techniques, and incorporate findings into defensive systems.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
7+ years of experience in applied machine learning, AI infrastructure, or safety-critical systems.
3+ years of experience in a senior, staff, principal, or equivalent technical leadership role.
Deep understanding of transformer-based architectures and large language models.
Experience developing, implementing, or evaluating safety interventions for LLMs.
Strong knowledge of adversarial machine learning, misuse prevention, and model robustness techniques.
Experience identifying and mitigating adversarial behaviors, edge-case failures, and model misuse scenarios.
Demonstrated ability to define long-term technical strategy, influence cross-functional stakeholders, and drive organizational technical direction.
Strong written and verbal communication skills.
Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field.
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.