AI Engineer
About the role
This position is responsible for designing, building, and maintaining the AI systems: the LangGraph agents, retrieval-augmented generation (RAG), multi-provider LLM orchestration, and prompt/evaluation machinery that turn customer conversations and uploaded documents into grounded, auditable cost estimates. Good leadership skills and the ability to work in US time zones as required are both prerequisites.
What you'll do
- LLM Agents & Orchestration
- Build and maintain LangGraph agents: ReAct retrieval sub-agents, custom middleware (call-limit, completion signaling), Pydantic-typed graph state, and streaming flows.
- Metrics / Measure of Success: Agent task success rate.
- Retrieval-Augmented Generation
- Own retrieval quality end to end: agentic retrieval with collection routing, LLM-driven metadata filters and similarity thresholds, query rewriting, semantic chunking, and citation tracking.
- Metrics / Measure of Success: Retrieval precision/recall on the eval set; grounding/citation accuracy. Groundedness score on the standard eval set.
- LLM Provider Integration
- Providers behind the abstract interface: AWS Bedrock, OpenAI, self-hosted vLLM.
- Backend Services & Data
- The platform's API services, relational and domain-knowledge data stores, the vector store, and observability/tracing for AI workflows.
- Technical Leadership & Collaboration
- Set technical direction for the AI workstream; review designs and code, mentor engineers, and align with product stakeholders. Break the status quo.
- Metrics / Measure of Success: Delivery against roadmap; quality of reviews and design decisions; growth and unblock rate of mentored engineers; stakeholder satisfaction.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, Data Science, or Mathematics.
- Minimum of 6 years of progressive software engineering experience, including 1+ years building production LLM-powered or machine-learning systems (agents, RAG, prompt/evaluation pipelines).
- Proficiency in Python 3, with strong knowledge of modern Python best practices, asynchronous programming, testing, and package management.
- Proven delivery in a fast-paced environment and availability to overlap U.S./Eastern Time business hours for client meetings.
- Expert in production software engineering and in building LLM-powered systems: agents/orchestration, retrieval-augmented generation, and prompt/evaluation pipelines; fluent in using AI development tools to accelerate delivery.
- Deep understanding of applied AI engineering: retrieval and grounding, embeddings and vector search, multi-model integration with cost/latency tuning, and evaluation-driven iteration. Working knowledge of the cost-estimation / Work Breakdown Structure domain, or the ability to learn a deep domain quickly.
- Ability to analyse complex, non-deterministic AI behaviour from real production traces and deliver scalable, well-tested solutions fast under time pressure (e.g., diagnosing over-retrieval, hallucination, or estimate/sizing collapse and landing a measurable fix).
- Exceptional ability to communicate technical and model-behaviour concepts clearly to non-technical and client stakeholders; confident in live U.S./ET client meetings and disciplined in clear, async written updates across time zones.
- Proven ability to make critical, time-sensitive decisions with limited information: balancing accuracy, latency, cost, and risk; high ownership and motivation, with a consistent bias to action and willingness to go the extra mile to deliver.
AI Alerts shares third-party job opportunities for informational purposes only. We are not the employer and are not involved in the hiring process. Always verify the company and role through official channels before applying, and never pay to apply, train, onboard, process documents, or secure a job offer. Legitimate employers do not ask applicants for money. Read our Terms to learn more.