HamburgerMenu
hirist

Prestige Trucking Insurance - Founding Engineer - Architecture & Reliability

Prestige Trucking Insurance
8 - 15 Years
Multiple Locations

Posted on: 02/09/2026

Job Description

Founding Engineer - Architecture & Reliability

Prestige Trucking Insurance is a US commercial trucking insurance agency, hiring in India. We've built Insure Stack: an agent uploads an intake form and it quotes every major trucking carrier in parallel, targeting all carriers returned in under 10 seconds. The product is live with real users. It works - and it has hard reliability problems that are the reason this role exists.

The role :

You'd be the senior-most engineer on a small team, owning architecture, reliability, and technical direction - and writing code every week. This is a builder's seat first and a leadership seat second. As we grow, you'd help shape and hire the team beneath you.

The honest version of the problem :

Our target is all carriers quoted in parallel, under 10 seconds. Today, when a quote fails, a human spends 1 - 2 hours recovering it by hand - and those manual fixes often introduce errors and rarely make it back into the automation, so the same failures recur. We believe the gap is architectural, not a matter of effort.

We're looking for someone who reads that and immediately starts forming hypotheses: sequential calls where there should be parallel fan-out, browser sessions that cold-start on every request, missing per-carrier timeout budgets, no partial-result path, no trace artifacts so every failure gets re-diagnosed from scratch, manual fixes never fed back into the system. You'd work out which of those it actually is, and rebuild it.

The system you'd inherit :

Next.js / React / TypeScript on Vercel; Python on Google Cloud Run (including a Cloud Run Job that runs our Playwright carrier automation); document extraction on GCS; a commission workflow integrated with IVANS; Auth0, Stripe, Sentry. You're not starting from scratch - you're taking over a live system and making it reliable.

What we need :

- 8+ years building production systems, with real architecture ownership at senior / staff / principal / lead level

- Deep Python - you've made async Python behave correctly under genuine concurrency

- You've owned a system where latency was the requirement, and can speak specifically to how you found and removed the time

- Reliability engineering: timeouts, retries with backoff, circuit breakers, idempotency, graceful degradation, observability that surfaces root cause

- You can take over an existing production system and improve it without a rewrite

- You've worked somewhere small enough that you owned outcomes, not just tickets

- Clear written communication; the team is distributed

Strong pluses :

- GCP depth (Cloud Run, GCS)

- Experience with browser automation or systems that depend on third-party interfaces you don't control

- Regulated-domain background (insurance, fintech, healthcare)

- Having inherited a struggling system and made it reliable

What this is not :

Not a management-only role - you'll be hands-on for the foreseeable future. And not distributed-systems-at-hyperscale: our problem isn't millions of requests per second, it's a handful of third-party websites that change without warning and must be handled correctly every time.

Working hours :

Starting hours are 9:00 AM EST for the first 90 days while you get situated with the platform and the team. After that, there may be room to revisit the schedule.

Our hiring process :

- A technical interview with our team

- A mock trial run - we give you a real problem from our platform and you work through it, so we can both see whether it's a good fit

- We're upfront about the mock trial because it's how we hire. We care more about what someone can do with a real problem in front of them than how they interview.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...