About
I work best where research has to turn into something real.
Most of my work lives somewhere between AI engineering, scientific ML, and the practical question of whether a system is actually useful once it leaves the slide deck.
Overview
I'm Ruman Shaikh. I work across AI engineering, scientific machine learning, and computational biology, and I'm currently doing an MSc in Biomedical Engineering at Imperial College London with a focus on computational bioengineering.
At IBM Client Engineering, I worked between IBM Research and client delivery. A lot of the job was taking ideas straight from research teams, figuring out if they actually made sense for a client problem, and then building the validation, benchmarks, and engineering around them. That work covered everything from RAG and fine-tuning to multi-agent systems, NL-to-SQL, climate risk modelling, and post-training source attribution work for LLMs. Before that, at Oracle Health, I worked on CMS-aligned healthcare data pipelines, analytics workflows, test automation, and performance engineering on large patient datasets.
I like work that sits between ideas and execution. I'm usually less interested in whether something can be demoed once, and more interested in whether it holds up, how it should be measured, and whether it survives real constraints. Long term, I want to keep working on AI and computational methods that are actually useful in science, health, and other high-impact settings.
Current Research
My dissertation at Imperial is on physics-informed operator learning for microbial dynamics, supervised by Prof. Reiko Tanaka. The core question is whether a generalised model can infer hidden biological parameters better than retraining a separate physics-informed model for every dataset. The work sits in the generalised Lotka-Volterra setting and compares baseline PINNs, PI-DION, and an exploratory QPINN setup.
More broadly, I'm interested in biological priors, neural scaling behaviour, reinforcement learning, operator learning, and the kinds of modelling choices that matter when data is scarce and the system itself is messy.
Tools And Areas
Languages & Libraries
- C
- C++
- Python
- Java
- SQL
- LLMs
- HTML
- CSS
- React.js
- LangChain
- LlamaIndex
- PyTorch
- TensorFlow
- scikit-learn
- OpenCV
- matplotlib
- CrewAI
- NLTK
- Hadoop
- Spark
- Hive
- Livy
- CUDA
Areas & Platforms
- Data Science
- Data Engineering
- MCP
- AI Agents
- Multi-agent Orchestration
- Automation
- LangGraph
- Machine Learning
- RAG
- Vector Databases
- Streamlit
- IBM Code Engine
- Docker
- Git
- GitHub
- MLOps
- Signal Processing
- GPU Programming