Design and optimize production AI systems, focusing on inference, model serving, and performance engineering for freight and supply chain applications. Profile bottlenecks in GPU, CPU, memory, and network resources to improve throughput, latency, and cost efficiency. Demonstrate exceptional technical depth by designing and executing rigorous benchmarks and experiments to drive engineering decisions.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
“I’ve seen things you people wouldn’t believe.”
— Blade Runner
Now we’d like to see what you’ve built.
We’re recruiting a Systems & Research Engineer for a fast-growing US applied AI company building production systems for freight and global supply chains.
Founded by engineers from MIT and Stanford, the company raised $6M in seed funding in early2026 and is already operating AI products at serious real-world scale.
One of its core fraud and identity platforms now screens approximately 3,000 drivers every day.
The team is small.
The ambition isn’t.
And the engineering bar is deliberately high.
What You’ll Actually Do
This is not another generic AI Engineer role.
You’ll sit at the intersection of:
AI systems.
Research.
Inference.
Model serving.
Performance engineering.
Distributed systems.
Evaluation.
Your job is to understand how production AI systems actually behave.
Where is the bottleneck?
GPU?
CPU?
Memory?
Network?
Serving architecture?
Concurrency?
Model choice?
You’ll form hypotheses, build benchmarks, test alternatives and use the results to make real engineering decisions.
A recent example involved benchmarking speech-to-text approaches and building a hybrid open-source and production system that outperformed vendor alternatives on both quality and economics.
That’s the level of problem we’re talking about.
We Want the Experiment, Not Just the Percentage
A resume saying:
“Reduced inference latency by 37%.”
isn’t enough.
We want to know:
What was the baseline?
Interested in remote work opportunities in Development & Programming? Discover Development & Programming Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
What did you think was happening?
How did you test it?
What alternatives did you benchmark?
What did the data show?
And most importantly:
What engineering decision changed because of the experiment?
You need to be able to walk us through at least one serious benchmark or experiment you personally designed and ran involving areas such as:
Model serving
Inference
Evaluation
Retrieval
Speech systems
Agent infrastructure
The strongest candidates think like researchers but ship like engineers.
The role specifically requires performance engineering on AI/model systems rather than distributed systems with no model in the loop.
The Kind of Work You Could Be Doing
Profiling production AI systems and identifying GPU, CPU, memory or network bottlenecks.
Benchmarking serving frameworks such as vLLM, SGLang and alternative architectures.
Investigating inference throughput and latency.
Evaluating open-source versus closed-source models.
Optimizing workloads across cost, quality and concurrency.
Building evaluation infrastructure.
Understanding model behavior under real production traffic.
Reasoning about distributed systems where there is genuinely a model in the loop.
Turning research findings into production architecture decisions.
And occasionally proving that everyone’s first assumption was wrong.
Who Could Be Right?
Your current title might be:
Research Engineer
ML Systems Engineer
AI Infrastructure Engineer
Inference Engineer
Performance Engineer
ML Platform Engineer
Systems Engineer
Experience inside sophisticated ML infrastructure environments is highly relevant.
Think engineering problems similar to those encountered at Google DeepMind, Meta, Stripe, Amazon AGI, Anthropic, Together AI, Fireworks AI, Baseten or similarly strong AI organizations.
But pedigree alone won’t get you through.
We’re looking for evidence of exceptional technical work.
The strongest candidates can explain their experience like this:
Here was the hypothesis.
Here was the baseline.
Here was the benchmark.
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Here’s what we discovered.
Here’s what we changed.
What This Role Is NOT
This is not DevOps.
It isn’t frontend.
It isn’t conventional full-stack development.
It isn’t Web3.
And pure distributed-systems experience, however impressive, isn’t enough if there has been no meaningful AI or model component.
There needs to be a model in the loop.
AI Coding Agents
This team uses coding agents heavily.
Every day.
The philosophy is simple: excellent engineers should increasingly spend their time deciding what should be built, how it should work and whether the result is correct, rather than manually producing every line.
Your ability to work effectively with modern coding agents will be assessed during the interview process.
Remote Really Means Remote
This role is 100% remote across the USA.
New York.
San Francisco.
Austin.
Seattle.
Boston.
Miami.
Denver.
Wherever you do your best work.
The company operates asynchronously, with no fixed working-hour or timezone-overlap requirement. There is a San Francisco office available if you want it, but attendance is not required.
Compensation
USD $220K–$300K Total Compensation + Equity
USD $300K is the ceiling.
This is high-bar, opportunistic hiring rather than a volume recruitment campaign. The company is building an ongoing pipeline and wants exceptional engineers, not simply more engineers.
The Bar
One of the founders has a very simple hiring philosophy:
He wants engineers joining the company who make the existing engineering team better.
That means the bar is high.
Deliberately.
If you’ve done genuinely exceptional work around AI systems, inference, model serving, evaluation or performance engineering, we want to hear from you.
And when you apply, don’t just tell us what you built.
Tell us about the experiment.
Systems & Research Engineer | Applied AI | USD $220K–$300K + Equity
100% Remote Across the USA | Fully Async | San Francisco Office Optional
Model Serving | Inference | ML Systems | Performance Engineering
Similar Jobs
Explore other opportunities that match your interests