Founding AI Engineer (Vision-language models)
Spotted 1h agofulltime
Job description
About this role
Employer-provided description, formatted for easier reading.
We're recruiting a Founding AI Engineer for an early-stage SF startup building AI smart glasses for field teams. Think technicians in data centers, energy and aerospace getting a co-pilot that sees what they see and walks them through the job.
You'd own the AI core: vision-language models running on real hardware, out in the field.
What you'd own:
- A production agentic VLM pipeline on the glasses: multi-step, tool-using visual reasoning against real inspections and procedures
- The eval and data flywheel: eval harnesses, failure capture, and customer data feeding fine-tunes that make each release better
- Edge inference and model orchestration that holds up when connectivity and latency get rough
- Real-time voice and video interfaces for the glasses
- Fine-tuning open models (SFT, RLHF, quantization) for on-prem deployments
Stack:
Python, PyTorch, Hugging Face, vLLM, Triton, TensorRT, ONNX.
You're probably a fit if you have:
- Hands-on training, fine-tuning or post-training of a vision-language or video model. Not just prompting or calling an API.
- Done that work at a self-driving, robotics, AR, top AI lab or funded AI startup
- About a year or more of that work, or a paper at a top venue
- Built agent loops with real tool use, not just RAG or a chatbot
- You're based in the US and can be in the SF office 5 days a week
Who this isn't for:
- IT services, consulting or non-tech enterprise backgrounds
- A run of short startup stints under a year
- Pure researchers who've never shipped to real users
- Language-only LLM folks with no vision work
- Anyone outside the US right now
Pay:
$180K to $230K base plus 0.25% to 0.75% equity.
Visa
OPT or H-1B transfers are fine. They can't sponsor a new visa.
This role is recruited by Jack. Apply and we'll set up a quick 15 minute call.
Interested in this role?Continue on Jack's careers page.
Apply on Jack