Tahasanul Ibrahim · Essen, Germany

Machine Learning Engineer.
Data Scientist.

MLOps · Forecasting · Industrial AI

I build forecasting and fault detection systems in Python, run them in production and work on AWS, with six years of experience in manufacturing and research. My work covers the full lifecycle, from PyTorch modelling and AWS workflows to monitoring and industrial integration.

Current focusForecasting and MLOps
Core stackPython · AWS · PyTorch
Industrial integrationOPC UA · MQTT · REST APIs
DeliveryEnd to end ML lifecycle

Capability map

Connected skills, grounded in real work.

The range is broad and the thread is consistent. I turn uncertain operational data into models, interfaces and decisions people can use.

01

Quantitative ML

Forecasting, optimisation, regression, classification, anomaly detection, time-series modelling, and model comparison for energy and industrial decisions.

ForecastingOptimisationPython
02

Industrial AI

Production quality control, adaptive fault detection, expert-knowledge fusion, maintenance guidance, real-time monitoring, and decision dashboards.

Evidence TheoryEnsemblesFMEA
03

Computer vision research

Evidence-fused object-detection ensembles and uncertainty-guided training, with reproducible GPU experiments and clearly bounded research evidence.

Faster R-CNNCUDAOptuna
04

Cloud and data systems

Reproducible ML workflows, time-series data modelling, AWS services, SQL and NoSQL stores, Databricks analytics, and API-connected services.

SageMakerTimestreamDatabricks
05

Engineering and MLOps

Dockerised services, CI/CD, Linux, Git, FastAPI, REST APIs, monitoring, self-hosted infrastructure, guarded configuration changes, and secure secrets handling.

DockerCI/CDFastAPI
06

Delivery and translation

Requirements discovery, technical specifications, project planning, partner communication, expert-knowledge translation, and early-stage scope-to-prototype work.

Project planningDocumentationPartner delivery

Independent projects

Tools and infrastructure I build and run.

Personal engineering work that runs alongside my professional roles, in applied inference, CI/CD and infrastructure. Each entry distinguishes a released tool, a tested prototype or a personally operated system. Expand an entry to see my contribution and its deployment context.

01Network configuration control planePersonal infrastructure · 07/2026 - Present+
  • Built configuration tooling for a personally operated RouterOS and WireGuard network, with validated inventory, an offline compiler and encrypted backups.
  • Implemented a guarded write path with a snapshot, automatic rollback protection, explicit confirmation and post change verification.
  • Designed drift reconciliation around the live network state, keeping CI read only and configuration changes under operator control.
Evidence and scope
  • Five automation managed RouterOS devices
  • Network operation evidenced since 05/2023
  • Personally operated infrastructure
  • Private implementation; no configuration files published
02ModelDeck, AI service telemetryPersonal project · Packaged integration+
  • Built a bridge that collects AI service usage and quota metrics from three providers and publishes sensors through MQTT discovery.
  • Implemented account authentication, token refresh and scheduled collection, with a FastAPI backend and React interface.
  • Delivered stable and nightly add-on packaging with automated release synchronisation, a 97% coverage gate and message broker integration tests.
Evidence and scope
  • Python, MQTT and container delivery
  • Tested provider integration and packaging
  • Provider and message broker integration tested
03Declarative service deploymentPersonal engineering project · Working v1; rewrite in progress+
  • Built a catalogue driven tool that renders container configurations from declarative service definitions and installation profiles.
  • Designed explicit conflict checks for service names, ports and data paths before applying a deployment plan.
  • Developed a phased engine rewrite with documented acceptance criteria, automated tests and validation against a disposable Linux runtime.
Evidence and scope
  • Working v1; v2 engine rewrite in progress
  • Python, FastAPI and Pydantic
  • Validated in a sandbox, not a production deployment
  • Private implementation
04Private cloud infrastructure and automation labPersonal infrastructure · 03/2021 - Present+
  • Designed and operated a Proxmox private cloud with Linux virtual machines and containerised services.
  • Built and scaled a Debian k3s cluster to five nodes, providing ephemeral self hosted CI runner pods.
  • Used the lab to develop and test application delivery, automation and operational tooling.
Technical scope
  • Proxmox and Linux
  • Docker and k3s
  • Self hosted CI infrastructure

How I work

Technical depth, shared understanding.

I work closely with the people who use a system, from agreeing on the problem to supporting delivery.

Understand the problem

Translate operational needs into clear requirements, a realistic scope, and a first prototype that helps the team make decisions.

Connect data and expertise

Bring together machine data, process knowledge, and FMEA context so model outputs make sense to the people acting on them.

Follow through on delivery

Connect models to APIs, dashboards, cloud services, and industrial protocols, with monitoring and documentation for ongoing use.

Work across disciplines

Coordinate technical work with manufacturing teams, researchers, and project partners. Make progress and tradeoffs clear, and support students and junior colleagues.

Computer vision research

Understanding model uncertainty.

I study how Evidence Theory can combine object detection models and guide training, using uncertainty-aware ensembling and uncertainty-guided loss weighting on reproducible GPU experiments. The work is available as a public preprint and has not been peer reviewed.

Read the public preprint
Research scopePreprint, not peer reviewed. A method study on PASCAL VOC 2012, with results under revision and broader generalisation still open.
Evidence fusionDempster-Shafer combination of multiple detectors in place of non-maximum suppression
Uncertainty-guided trainingDynamic loss weighting driven by per-prediction uncertainty rather than a fixed schedule
Reproducible experimentsCUDA training runs with Optuna search under a fixed evaluation protocol
Bounded scopeA PASCAL VOC 2012 method study, with reported results under revision and not quoted here
PyTorch / torchvisionFaster R-CNNSwappable CNN backbonesCUDA DataParallelOptunamAP · F1 · precision · recallDS conflict K and certainty Phi

Professional experience

From research to production machine learning.

A full timeline of roles and project work, in the same order as my CV.

10/2023 - Present

Machine Learning Engineer (Doctoral Project) Promotionskolleg NRW · Bochum, Germany

  • Reduced a 220 GPU-hour model evaluation workload to 3.7 GPU-hours using a two-tier GPU and CPU cache, reusing model outputs for repeated evaluation instead of rerunning GPU inference.
  • Implemented 2-rank distributed training with PyTorch, NVIDIA Collective Communications Library (NCCL) and activation checkpointing, fitting model training within 16 GB GPU memory while preserving identical gradients.
  • Built a Python orchestration layer for seven training chains in quota-capped GPU environments, using rolling quota accounting, checkpoint-exact resume and SHA-256 dataset handoff for unattended execution.
  • Developed uncertainty-aware Dempster-Shafer ensembles for object detection and benchmarked them against non-maximum suppression (NMS) and voting baselines, backed by a 213-test suite that regenerates every reported result.
  • Established LLM-assisted code review and quality-check workflows, with defined acceptance criteria and a deterministic scanner that blocks non-compliant output.
PyTorchMLOpsDistributed TrainingCUDAOptunaModel EvaluationUncertainty QuantificationLLM Agents
06/2026 - Present

Media Classification and OCR Service Independent project

  • Built a media classification and OCR service using batched ONNX Runtime inference with a ViT-base classifier, MobileNetV2 tagging and PP-OCRv5 text recognition.
  • Pinned every model by SHA256 with recorded licence acceptance in a versioned catalogue, behind an enforced 90% test coverage gate.
ONNX RuntimeVision TransformersMobileNetV2PaddleOCRAES-256-GCMpytest
09/2025 - 01/2026

Data Scientist be.storaged GmbH · Oldenburg, Germany

  • Prepared demand-forecasting features from weather, time of day, holidays and historical consumption, supporting model evaluation and iteration.
  • Developed and validated power-demand forecasting models, benchmarking machine-learning candidates against classical statistical approaches.
  • Trained CUDA-accelerated transformer models for forecasting and handed evaluated model candidates to the team.
  • Supported model monitoring and retraining in AWS workflows spanning Lambda, SageMaker, Timestream, DynamoDB and Grafana.
AWSSageMakerLambdaTimestreamDynamoDBGrafanaCUDATransformers
05/2025 - Present

Organisation-Wide CI/CD and Release Platform Independent project

  • Built a reusable GitHub Actions platform serving 14 repositories with shared CI, promotion, nightly and release workflows consumed by reference.
  • Replaced access tokens with a least-privilege GitHub App identity to publish PyPI and GHCR releases behind Semgrep and enforced coverage gates.
Reusable WorkflowsARCk3sGitHub AppsSemgrepTrivy
01/2020 - 07/2025

Machine Learning Engineer, Research Associate South Westphalia University of Applied Sciences · Soest, Germany

  • Lead technical planning and delivery across publicly funded industrial AI projects, coordinating manufacturing, quality and software stakeholders.
  • Built reusable Python workflows with multi-GPU CUDA ResNet training and Optuna searches distributed across a SQL-backed machine grid.
  • Built and operated retrieval-augmented generation (RAG) pipelines for internal users, with embedding-based retrieval over a vector store, multi-step orchestration and schema-defined tool calling.
  • Built the data pipelines and Databricks analytics behind the deployed models, released through GitLab CI/CD.
  • Developed sensor fusion and uncertainty quantification (UQ) for anomaly detection on live machine data, attaching an explicit uncertainty measure to every output.
PythonPyTorchCUDAMulti-GPUOptunaRAGGitLab CI/CDDatabricks
11/2021 - 12/2022

Data Scientist AI Information Fusion and Operator Guidance · ZIM-funded, with ShoutR Labs

  • Lead and developed automated rule discovery and Dempster-Shafer evidence fusion into ranked repair recommendations.
  • Integrated the operator-guidance workflow with frontend and production systems through REST APIs and OPC UA mappings, coordinating daily technical delivery.
Dempster-ShaferEvidence TheoryRule DiscoveryREST APIsOPC UAFMEAPython
01/2020 - 09/2021

Machine Learning Specialist WiTraPres, Knowledge Transfer and Adaptive Fault Detection · BMBF-funded, with GEDIA Automotive Group

  • Built adaptive fault-detection workflows with feature engineering on live OPC UA and MQTT data, scoring the streams every 15 seconds.
  • Ran the models on line-side industrial hardware rather than a back-end server, and validated them across machines, production lines and process recipes.
  • Delivered the surrounding system as Docker microservices with Django and Streamlit operator dashboards, published in peer-reviewed IEEE work.
PyTorchOPC UAMQTTDockerDjangoStreamlitGitLab CI/CDDatabricks
05/2019 - 12/2019

Systems Engineer (Working Student) | Master’s Thesis GEDIA Automotive Group · Attendorn, Germany

  • Built adaptive fault-detection models on live automotive production data using XGBoost, LightGBM, ensemble methods and Evidence Theory.
  • Selected features and tuned model combinations and hyperparameters through comparative experiments, using principal component analysis (PCA) for dimensionality reduction.
  • Co-developed a Python Kafka branch on a shared platform, linking consumed machine records to MySQL queries and the failure mode and effects analysis (FMEA) diagnostic workflow.
PythonXGBoostLightGBMEnsemble MethodsPCAEvidence TheoryKafkaFMEA
01/2019 - 09/2019

IoT Data Engineer IoT Air Emission Monitoring and Forecasting · Universidad de El Salvador

  • Managed remote operations and troubleshooting for six IoT emission-monitoring stations using distributed devices and live sensor data.
  • Built the end-to-end MQTT flow connecting the stations, a long short-term memory (LSTM) forecast and a live dashboard with bidirectional commands for a SaaS cloud application.
MQTTLSTMTime-Series ForecastingIoTSaaS CloudDashboards

Technical expertise

Tools for the full ML lifecycle.

From experimentation and data integration to deployment, monitoring, and infrastructure.

Machine learning and statistics

Time-series forecasting (LSTM, 1D-CNN) · anomaly detection · classification · computer vision · object detection · incremental learning · transformers · CNNs · ensemble methods · Bayesian methods · K-means clustering · PCA · feature engineering · Evidence Theory

Programming and ML libraries

Python · C++ · Java · JavaScript · Bash · PyTorch · torchvision · scikit-learn · ONNX Runtime · Pandas · NumPy · OpenCV · XGBoost · LightGBM · Optuna · SQLAlchemy · Pydantic · pytest

MLOps and cloud

AWS SageMaker · Lambda · S3 · Athena · Timestream · DynamoDB · Step Functions · SQS · Docker · Kubernetes (k3s) · Terraform · MLflow · Weights & Biases · Grafana · GitLab CI/CD · GitHub Actions

Data and industrial integration

Databricks · SQL and NoSQL · Kafka · MQTT · OPC UA · REST APIs · ETL/ELT · MySQL · Oracle

Applications and infrastructure

FastAPI · Flask · Django · Streamlit · Power BI · React · Linux · k3s · Proxmox · Traefik · self-hosted CI runners

Inference and GenAI

Batched ONNX Runtime inference · self-hosted LLM runtime deployment · LLM provider API integration · multi-agent development workflows

Publications

Industrial AI and computer vision research.

Education and languages

Engineering foundation, research context.

Dr.-Ing. Machine Learning (ongoing)Promotionskolleg NRW, Bochum, Germany · 10/2023 - Present
M.Sc. Systems Engineering and Engineering ManagementSouth Westphalia University of Applied Sciences · 10/2015 - 12/2019
B.Sc. Electrical and Electronic EngineeringAmerican International University, Dhaka · 09/2009 - 02/2014

Languages

GermanB2 (CEFR)
EnglishC2 (CEFR)
BengaliNative

Let’s connect

Working on a data or machine learning challenge?

I am based in Essen and looking for Machine Learning Engineer, Data Scientist and MLOps opportunities, primarily in Germany and also across the EU. I would be glad to discuss your team’s work and where my experience could help.