I build AI systems where quality is measured, not assumed.
Ten years across document intelligence, healthcare, video analytics, and financial services. The through-line is evaluation-driven development — defining quality metrics up front, building the benchmarks and regression suites that hold a system to them, and running offline and online evaluation to decide what actually ships.
Most recently at Apple, I designed an intelligence layer for enterprise document analysis on Gemini 2.5 and multi-agent systems, with acceptance criteria expressed as Pydantic schemas so non-conforming output is rejected rather than discovered later. Before that, healthcare and video models serving law enforcement agencies in Europe and North America.
Selected results
Extraction accuracy, critical metadata fieldsApple
98%+
Token classification F1, PHI detectionBeeKeeperAI
97.3%
False negatives, real-time transaction monitoringAccess Bank
−46%
Pipeline time-to-market, enterprise integrationsTribes.ai
−45%
Customer satisfaction, sentiment analysis modelsAccess Bank
+20%
measured score
change vs. baseline
scale 0–100%
Experience
Senior AI Engineer — Document Intelligence
Apple (Contractor)
Feb 2025 – Aug 2026
- Evaluation-driven development. Reached 98%+ accuracy on critical metadata extraction by expressing acceptance criteria as Pydantic schemas — non-conforming output is rejected at the boundary, which makes quality measurable rather than a matter of opinion.
- Agentic AI and RAG. Designed a first-of-its-kind intelligence layer for enterprise document analysis on Gemini 2.5 and multi-agent systems with LangGraph, blending agentic reasoning with retrieval.
- Tool use and orchestration. Built multi-agent orchestration over the Model Context Protocol for tool-driven task execution across enterprise repositories.
- Technical direction. Partnered with product managers and domain stakeholders to turn ambiguous needs into designs, and drove technical decisions through working demos and iterative feedback.
Senior ML Engineer, Platform Lead
BeeKeeperAI & RigrAI
Jan 2023 – Feb 2025
- Model evaluation and quality. Delivered BERT-based classification at 97.3% F1, with offline benchmarking and monitoring that caught quality regressions before release.
- Deployment, scaling, observability. Engineered high-throughput serving with FastAPI, Docker, and Nvidia Triton across GCP and Azure, tuned for latency, throughput, and cost against global SLO and SLA targets.
- LLM optimization. Fine-tuned Llama models with LoRA and weight orthogonalization for domain-specific summarization, and implemented prompt-injection defenses and validation guardrails.
- Mentoring. Served as technical advisor to enterprise customers and guided engineers on architecture, evaluation, and production trade-offs.
Senior Data Engineering Lead
Tribes.ai (B2B SaaS)
Jan 2022 – Nov 2022
- Online evaluation and experimentation. Used A/B testing to analyze model behavior and refine behavioral analytics models, maximizing enterprise customer ROI.
- Regression suites. Implemented automated data quality checks on Kubernetes and GCP, catching regressions in pipeline output before they reached consumers.
- Pipelines at scale. Designed automated pipeline templates that cut time-to-market by 45% across enterprise data integrations.
Enterprise Data Science Engineer
Sterling Bank PLC
Sep 2021 – Jan 2022
- Production ML systems. Designed end-to-end models integrating disparate enterprise data sources and APIs.
- Large-scale data processing. Architected workflows on Databricks and Azure Synapse, and built services in Golang and Python.
Data Science Engineer
Access Bank PLC
Mar 2017 – Sep 2021
- Error analysis at scale. Reduced the false-negative rate by 46% through systematic error analysis over real-time monitoring of millions of daily records with Spark Streaming and Kafka.
- Applied NLP. Built and deployed sentiment analysis models over customer feedback, improving customer satisfaction metrics by 20%.
Capabilities
Evaluation & quality
- Evaluation-driven development
- Offline & online evaluation
- Regression suites
- LLM-as-a-judge
- Benchmarks
- Quality metrics
- Error analysis
- Human evaluation
- RAGAS
- DeepEval
- Acceptance criteria
- Model monitoring
Search, retrieval & agents
- RAG
- Tool-using agents
- LangGraph
- MCP
- Vector databases
- Embeddings
- Semantic retrieval
- Re-ranking
- LangChain
- Knowledge graphs
- Function calling
- Prompt engineering
- ChromaDB
- FAISS
- Neo4j
Models & training
- PyTorch
- Fine-tuning
- Gemini
- Llama
- TensorFlow
- scikit-learn
- Hugging Face
- LoRA
- PEFT
- GPT-4
- Claude
- NLP
- Deep learning
- Classical ML
Platform & infrastructure
- Model serving
- Observability
- Docker
- Kubernetes
- Nvidia Triton
- FastAPI
- CI/CD
- GitHub Actions
- Jenkins
- Airflow
- Kafka
- Responsible AI
Languages, data & cloud
- Python
- SQL
- Spark
- AWS
- Databricks
- Go
- TypeScript
- JavaScript
- PySpark
- Spark Streaming
- SageMaker
- Redshift
- EKS
- S3
- GCP
- Vertex AI
- Azure AI Foundry
- PostgreSQL
- Snowflake
Publications
Reducing Hyperparameter Tuning Costs in ML, Vision and Language Model Training Pipelines via Memoization-Awareness
A. Essofi, R. Salahuddeen, et al.
Real-time Pothole Anomaly Detection: Toward Safer Roads in Developing Nations
Engineering Proceedings
Deep Learning Model for Potholes Data Augmentation
IEEE Conference
Full publication record on Google Scholar
Education & certifications
Mohamed bin Zayed University of AI
M.Sc. Machine Learning — full scholarship
Abu Dhabi, UAE · 2022–2024
Georgia State University
M.Sc. Statistics and Computer Science
Atlanta, USA · 2023–2025
Ladoke Akintola University of Technology
B.Tech Electrical and Electronics Engineering
2009–2014
- Microsoft Certified: Azure AI Engineer Associate
- Microsoft Certified: Azure Data Scientist Associate
- Microsoft Certified: Azure Data Engineer Associate
- Udacity: Deep Learning Nanodegree