Portrait of Xianglong Fu
Xianglong Fu
Data Scientist & ML Engineer
Master of Data Science, University of Melbourne
// BIOGRAPHY

About me

I'm a data scientist and ML engineer who turns ideas into working systems. I'm trained in statistics and machine learning, and I care about how models actually work under the hood — not just how to call them from a library.

I build end to end: data pipelines, statistical and ML models, and the interfaces around them. I use AI-assisted development to ship faster, always with careful human review and testing to keep every detail honest. I'm completing my Master of Data Science at the University of Melbourne.

Interests

  • ● Machine Learning
  • ● Deep Learning
  • ● Statistical Modelling
  • ● Data Engineering
  • ● Cloud & Distributed Systems
  • ● AI Applications (RAG)
  • ● AI Planning

Education

🎓
Master of Data Science
University of Melbourne · Melbourne, Australia
2025 — 2026
Expected Dec 2026. Coursework across machine learning, statistical modelling, and high-performance computing.
🎓
B.Sc. in Applied Statistics
Western University · London, Ontario, Canada
2020 — 2024
Foundations in probability, statistical inference, and data analysis.
// PROJECTS

Projects

Preprint · Under Review Testing hypotheses about correlations between brain activation patterns

Testing hypotheses about correlations between brain activation patterns

Second author · independent simulation experiments

  • Independently ran the simulation experiments for this statistical-methodology paper (second author), now a preprint on bioRxiv and under review at Imaging Neuroscience.
  • The experiments give the paper its empirical grounding — testing how the method performs under controlled, reproducible conditions.
Internship Internal scheduling & document tool for a medical clinic

Internal scheduling & document tool for a medical clinic

Google Maps API · RAG retrieval · AI-assisted frontend

  • Built internal operations tooling for a medical clinic: I implemented the Google Maps API route/distance planning and the retrieval-augmented (RAG) document search.
  • The distance planning helps staff plan visit routes more efficiently; the RAG search lets them ask questions over the clinic’s own documents instead of digging through files.
  • The frontend was delivered with AI assistance, under my review and testing.
Capstone Aroma & Taste — predicting molecular odor and taste

Aroma & Taste — predicting molecular odor and taste

PubChem · GoodScents · entity matching · modeling

  • Independently completed a full data-science pipeline to predict the odor and taste of chemical molecules from their structure.
  • Integrated and matched records across multiple public sources (PubChem and GoodScents), engineered features, and built predictive models.
  • Turned messy, multi-source inputs into a working model and a clear analysis — the whole journey from raw data to results.
Deep Learning LeNet-5 in pure NumPy

LeNet-5 in pure NumPy

NumPy · im2col + BLAS · Three.js · Hugging Face Spaces

  • A complete convolutional network — convolutions, pooling, backprop, and the full training loop — implemented in pure NumPy with no deep-learning framework. The point wasn’t to replace frameworks, but to understand exactly what they do under the hood.
  • Optimized the convolutions with im2col + BLAS for an ~11× speedup; the model reaches 96.3% accuracy on MNIST.
  • Packaged with an interactive Three.js visualization and deployed to Hugging Face Spaces so anyone can explore it in the browser.
Cloud · Distributed AusFlux — real-time data collection on Kubernetes

AusFlux — real-time data collection on Kubernetes

Kubernetes · real-time ingestion

  • Team cloud-computing project; I owned the data-collection layer, running real-time ingestion on Kubernetes.
  • Diagnosed and removed a rate-limiting bottleneck, increasing collected volume ~6× — the fix that kept the pipeline fed.
HPC Parallel big-data processing with MPI

Parallel big-data processing with MPI

mpi4py · distributed memory

  • Parallelized a large data-processing job with mpi4py across a distributed-memory cluster.
  • Handled 219 GB of data and achieved a 7.77× speedup over the serial baseline.
Work in Progress Go (围棋) AI — MCTS + self-play

Go (围棋) AI — MCTS + self-play

PyTorch · ResNet · MCTS · self-play

  • A reinforcement-learning project in the AlphaGo Zero style: a policy/value network trained through self-play, guided by Monte Carlo tree search (MCTS).
  • Work in progress — the search and self-play training loop is still being built out and validated.
Full-Stack · Web Minesweeper, a full-stack web game

Minesweeper, a full-stack web game

FastAPI · Vite · TypeScript

  • A full-stack Minesweeper web game — a faithful Windows-95-style take built on FastAPI, Vite, and TypeScript.
  • An AI solver is in early exploration: work in progress, currently learning the space before committing to any one approach.
// PUBLICATIONS

Publications