Lead ML/AI Platform Engineer
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead ML/AI Platform Engineer based in Ireland.
As a Lead ML/AI Platform Engineer, you will own the infrastructure that moves machine learning and AI from experimentation into reliable production systems. You will define technical direction for ML/AI architecture, tooling, standards, and platform strategy while partnering closely with data science, DevOps, backend engineering, product, and leadership teams. The role spans managed AWS machine learning services, open-source ML tooling, model serving, inference pipelines, and production integrations. You will also drive the engineering strategy for GenAI, LLMs, retrieval architectures, and agentic workflows. Your decisions will establish the foundation for future AI-powered products in a highly regulated financial technology environment. This is a fully remote independent contractor opportunity within a globally distributed team, with reliable overlap with US Pacific business hours.
Accountabilities
- Own the ML/AI platform: Lead the architecture and operation of training infrastructure, model serving, inference pipelines, model registries, feature pipelines, and production integrations.
- Set technical direction: Partner with the Data Platform Architect to define ML/AI architecture, engineering standards, tooling strategies, and build-versus-buy decisions.
- Build production ML systems: Oversee feature engineering, model training, tuning, registration, deployment, monitoring, and hosted inference using AWS managed services and open-source tooling.
- Drive GenAI and LLM engineering: Develop practical strategies for RAG, prompt engineering, evaluation, fine-tuning, model serving, and agentic workflows while balancing cost, latency, quality, and safety.
- Develop agentic capabilities: Evaluate and implement emerging agent technologies and workflow patterns, including MCP, stateless and stateful architectures, and appropriate guardrails for regulated environments.
- Own the serving layer: Design scalable ML services and API contracts that integrate cleanly with Java microservices, making informed decisions around performance, latency, throughput, and reliability.
- Enable research-to-production: Partner with Data Science to productionize models and create the infrastructure, pipelines, and deployment processes required to move experimentation into reliable production.
- Establish MLOps practices: Drive model monitoring, drift detection, reproducibility, experiment tracking, model registries, and cost observability in collaboration with DevOps.
- Shape the ML roadmap: Work with product, data, and engineering leadership to identify high-impact AI/ML opportunities and translate them into actionable technical roadmaps.
- Provide technical leadership: Mentor senior engineers, guide complex architectural decisions, and represent the AI/ML function in cross-functional discussions.
- Communicate technical strategy: Translate complex architectural trade-offs into clear design documents and executive-level recommendations for both technical and non-technical stakeholders.
Requirements:
- Senior engineering experience: 8+ years in software or ML engineering, including at least 5 years delivering production machine learning systems and owning complex, ambiguous problems end to end.
- Technical leadership: Demonstrated experience shaping ML strategy, mentoring senior engineers, and serving as a trusted decision-maker for challenging architectural problems.
- AWS ML expertise: Hands-on experience with Amazon SageMaker for training, tuning, hosted endpoints, and model registry, as well as Amazon Bedrock and AgentCore for GenAI and agentic applications.
- Open-source ML tooling: Practical experience with JupyterLab, Spark, MLflow, and related machine learning development and experimentation tools.
- AWS platform knowledge: Deep familiarity with services including S3, Athena, Redshift, Glue, Step Functions, and Lambda, combined with strong SQL skills for analytical and ML workloads.
- GenAI/LLM expertise: Production experience with RAG, prompt engineering, evaluation, and the practical trade-offs between quality, cost, latency, and safety.
- Vector search knowledge: Experience with vector databases such as pgvector or Pinecone, together with good judgment about when vector retrieval is appropriate versus alternative approaches.
- Programming skills: Deep Python expertise and strong knowledge of core ML libraries such as scikit-learn, pandas, NumPy, PyTorch and/or TensorFlow, and XGBoost or LightGBM.
- Java familiarity: Sufficient working knowledge of Java to review service code, define API contracts, and troubleshoot integrations with backend microservices.
- MLOps expertise: Strong understanding of model monitoring, drift detection, reproducibility, experiment tracking, model registries, deployment patterns, and cost observability.
- Cross-functional collaboration: Comfortable partnering with Data Science, DevOps, backend engineering, product, and leadership while maintaining clear ownership across shared technical boundaries.
- Communication: Excellent written and verbal communication skills, with the ability to produce both detailed technical designs and concise executive-level recommendations.
- Contracting requirements: Ability to work independently as an independent contractor through your own entity or an approved contracting arrangement, with reliable overlap with US Pacific business hours.
- Nice to have: Experience with ClickHouse or similar analytical databases, LLM fine-tuning techniques such as LoRA or QLoRA, streaming and real-time inference, Kafka or Kinesis, infrastructure-as-code, large-scale ML systems, or open-source ML contributions.
Benefits:
- Equity compensation package.
- Flexible Time Off (FTO) to take time away when needed to rest and recharge.
- Medical, dental, and vision coverage, with 100% of employee premiums covered where applicable.
- Disability and life insurance.
- Learning and career development opportunities within a growing technology environment.
- Remote-work setup reimbursement.
- Monthly phone and internet stipend.
- Team-building events, cultural activities, and company-wide gatherings.
- Paid time off for volunteering and community service.
- Half-day Fridays.
- 401(k) matching contribution.
- Opportunity to work on AI/ML infrastructure supporting products used by more than 1,600 financial institutions.
- Fully remote work environment as part of a globally distributed team.