Staff Applied AI Engineer, Agents & Model Adaptation
Senior individual-contributor role at an early-stage AI startup focusing on intelligent systems for fraud and risk detection. Key responsibilities include LLM research, model adaptation, agentic systems, and production AI infrastructure. Requires deep hands-on experience fine-tuning LLMs, building production AI agents, and developing evaluation frameworks.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Staff Applied AI Engineer, Agents & Model Adaptation
Compensation: $240,000-$275,000 + very strong equity
Location: New York, NY
Company: An early-stage AI startup building intelligent systems to identify complex fraud and risk across the insurance ecosystem.
We’re partnering with a venture-backed AI company using advanced data and machine learning to tackle sophisticated fraud across property and casualty insurance. By connecting signals across multiple parts of the ecosystem, the team is helping uncover complex networks, improve decision-making, and automate investigative workflows that have traditionally required significant manual effort.
The Role
This is a senior individual-contributor role for an engineer who wants to operate at the intersection of LLM research, model adaptation, agentic systems, and production AI infrastructure.
You’ll own key architectural decisions around which frontier and open-weight models should power different workflows, how those models are adapted to domain-specific tasks, and how agent performance is evaluated in production.
A major focus will be solving the trade-off between the reasoning quality of frontier models and the latency, cost, privacy, and scalability benefits of smaller open-weight models.
You’ll remain deeply hands-on while setting the technical direction for the wider AI engineering function.
What You’ll Do
- Own the base-model strategy across AI workflows, evaluating frontier and open-weight models and driving decisions around when to adopt, replace, or adapt them.
- Fine-tune open-weight LLMs using techniques such as LoRA, QLoRA, supervised fine-tuning, and preference optimization / DPO.
- Deploy adapted models as production-grade components within agentic workflows.
- Build a continuous data flywheel, transforming human corrections and production agent traces into training data, synthetic datasets, evaluation cases, and future fine-tuning runs.
- Develop evaluation frameworks that benchmark models and agents against realistic, domain-specific scenarios.
- Measure and optimize performance across accuracy, hallucination rates, false positives, cost, and latency.
- Design and productionize agent infrastructure covering tool use, context construction, orchestration, structured outputs, guardrails, and observability.
- Implement eval-gated deployment processes and full traceability across production agent runs.
- Own serving infrastructure for self-hosted models, including batching, quantization, latency optimization, and inference cost management.
- Red-team models and agent systems for issues including prompt injection, data leakage, and adversarial manipulation.
- Set a high technical bar through architecture reviews, mentorship, and clear written technical direction.
Searching for Development & Programming roles that provide visa sponsorship? Connect with international employers through Development & Programming Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
What You’ll Bring
- Experience across software engineering, applied machine learning, or related disciplines.
- Experience acting as a technical leader and making architectural decisions for production ML or LLM systems.
- Deep hands-on experience fine-tuning large language models, with production results you can discuss in detail.
- Experience building production AI agents that use external tools and generate structured outputs.
- Strong experience developing evaluation systems that meaningfully influence model or product decisions.
- Advanced proficiency with Python and PyTorch.
- Familiarity with modern LLM training and distributed-compute tooling.
- Comfort working in ambiguous environments where ground truth is imperfect and existing research does not provide an off-the-shelf solution.
Bonus experience includes:
- Shipping quantized or distilled models into latency- or cost-constrained environments.
- Diagnosing and addressing accuracy regressions caused by compression or serving optimizations.
- Entity resolution, graph ML, or classical ML for highly imbalanced detection problems.
- Agent benchmarking datasets and designing domain-specific evaluation suites.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Tech Stack
- Python
- PyTorch
- Hugging Face ecosystem
- LoRA / QLoRA
- Supervised fine-tuning
- DPO / preference optimization
- vLLM, TGI, Triton, or equivalent serving infrastructure
- DeepSpeed, Ray, or equivalent distributed training tooling
- Open-weight and frontier LLMs
- Agent orchestration and tool-use frameworks
- Evaluation and observability infrastructure
Why Join?
You’ll have unusually broad ownership across the AI stack from model selection and fine-tuning through agent architecture, evaluation, serving, and safety.
Rather than simply wrapping third-party APIs, you’ll be solving technically difficult questions around how reasoning-capable models can be adapted and deployed efficiently at significant production scale.
You’ll join an early-stage environment where technical decisions have immediate product impact, work directly on a challenging real-world problem, and help establish the foundations and engineering standards for the company’s AI platform.
The package includes company equity, 100% medical coverage plus dental and vision, and potential visa-transfer sponsorship for eligible candidates currently based in the U.S.
About People In AI
People In AI is a specialist recruitment partner connecting exceptional AI, machine learning, and data talent with some of the most ambitious technology companies in the market.
We work closely with candidates and businesses across the AI ecosystem, from early-stage startups to established technology organizations, helping build teams working on genuinely impactful problems.
Similar Jobs
Explore other opportunities that match your interests