Tariful Islam Tarif
AI/ML Engineer driving the design, deployment, and operation of production-grade machine learning, LLM, and data-driven systems. Expertly crafting sca...
About
AI/ML Engineer driving the design, deployment, and operation of production-grade machine learning, LLM, and data-driven systems. Expertly crafting scalable, low-latency inference pipelines, multi-agent LLM architectures, and end-to-end MLOps workflows with Python, FastAPI, Docker, and Kubernetes.
Skills
Programming Languages
Python
SQL
C++
Java
C
Machine Learning & Deep Learning
PyTorch
TensorFlow
Scikit-learn
ONNX Runtime
Model Training & Evaluation
Model Versioning
LLMs & Generative AI
LangChain
LangGraph
Retrieval-Augmented Generation (RAG)
Prompt Engineering
Multi-Agent Systems
LLM-as-Judge Evaluation
OpenAI API
Hugging Face
Data Science
Data Mining
Statistical Analysis
Exploratory Data Analysis
NL-to-SQL
Feature Engineering
Backend & APIs
FastAPI
REST API Design
Asynchronous Programming
Microservices Architecture
Databases
PostgreSQL
MongoDB
SQLite
Vector Databases (FAISS, Pinecone, Chroma)
MLOps & Observability
Prometheus
OpenTelemetry
Structured Logging
Drift Monitoring
Model Versioning
DevOps & Tools
Docker
Kubernetes
Git
GitHub
Canary Rollouts
Autoscaling
Projects
Intelligence Reliability Platform — LLM Observability & Evaluation
FastAPI
React
OpenAI API
Ollama
SQLAlchemy
Docker
Built a production-grade LLM observability platform featuring prompt versioning, risk controls, and multi-model fallback routing (OpenAI → Ollama → Mock) for high availability. Designed automated LLM evaluation pipelines using ensemble LLM-as-judge scoring across quality, safety, and hallucination-detection dimensions. Implemented real-time trace logging, cost tracking, latency metrics, and drift monitoring to support continuous model governance. Developed role-based access control (RBAC), audit logging, and secure inference workflows for enterprise-grade compliance.
Real-Time ML Inference System
FastAPI
ONNX Runtime
Redis
Prometheus
Docker
Kubernetes
Engineered a low-latency ML inference API supporting multi-model and multi-version serving. Achieved sub-100ms inference latency using ONNX Runtime and Redis-backed intelligent caching, cutting redundant compute by 80%+. Implemented Prometheus metrics, health checks, and OpenTelemetry-based structured logging for full-stack observability. Deployed via Docker Compose and Kubernetes with autoscaling and canary rollout support for zero-downtime releases.
Multi-Agent Medical Data Assistant
Python
Google Gemini
GPT-4o-mini
SQLite
NL-to-SQL
Built a multi-agent LLM system enabling natural-language querying of medical datasets (heart disease, cancer, diabetes). Implemented NL-to-SQL pipelines with automated routing between database query tools and medical web-search agents. Integrated LLM-based summarization to generate structured, evidence-based analytical insights for healthcare data.
Advanced Image Change Detection System
Python
OpenCV
ORB
SSIM
Designed a computer vision pipeline for robust before/after image comparison under camera misalignment. Implemented ORB feature matching, homography alignment, and SSIM-based structural change detection. Generated annotated outputs, heatmaps, binary masks, and structured JSON summaries to support auditability.
LangGraph-Style Multi-Agent Debate Workflow
Python
LangGraph
Graphviz
Built a deterministic multi-agent orchestration system with strict turn control and debate sequencing logic. Implemented memory nodes, semantic duplicate detection, and replayable JSON-based event logging. Visualized multi-agent workflows as DAGs to support debugging and evaluation.
Experience
Forward Deployed Engineer
INFINOZ
01/03/2026 - Present
Partnered directly with product and client teams as a Forward Deployment Engineer to design, build, and deploy applied AI solutions from research through to production. Architected and developed InfiChat, an AI-driven conversational assistant, owning model integration, system architecture, and end-to-end deployment. Led R&D on a real-time Voice Agent, integrating speech recognition (ASR), natural language understanding (NLU), and speech synthesis (TTS) into a unified conversational pipeline. Fine-tuned and evaluated Bangla Text-to-Speech (TTS) models, covering dataset preparation, model fine-tuning, and voice-quality evaluation to improve native Bengali speech synthesis. Collaborated cross-functionally with engineering and product stakeholders to convert experimental AI research into reliable, scalable, production-ready voice and chat-based AI components.
Education
B.Sc. in Data Science & Engineering
University of Frontier Technology Bangladesh (UFTB)
- Present
Relevant Coursework: Machine Learning, Deep Learning, Natural Language Processing, Data Mining, Statistics, Python Programming
Certifications
AI Engineer Bootcamp
Ostad