AI System Design2025Deep Learning

Brain tumor MRI classifier: in your browser

A privacy-first web app that sorts brain-MRI scans into four tumor types entirely on-device. The scan never leaves your machine, and a ~1.5 MB model returns an answer in about nine milliseconds.

Upload a scan, get a classification with live confidence scores, computed entirely in the browser.
Overview

Medical AI that keeps the data at home.

This is a web application that classifies a brain-MRI scan into one of four tumor types and shows real-time confidence for each. The twist: it runs the deep-learning model entirely in the browser using ONNX Runtime Web, so the image is analysed locally and never uploaded anywhere.

It's a study in doing more with less: a lightweight ~1.5 MB model delivering ~9 ms inference, deployed as a live app anyone can try.

Live demo

This isn't a recording. It's the real model.

The player above is a walkthrough. Below is the actual deployed app, embedded live: upload an MRI slice (or use a sample) and the same ~1.5 MB ONNX model classifies it in your browser, right now, with no server in between.

This is my own public app, or open it in a new tab instead ↗

Problem

Sensitive scans shouldn't have to travel.

Medical images are among the most private data there is. The usual pattern, uploading a scan to a server for a model to analyse, introduces privacy exposure, compliance overhead, network latency, and hosting cost. For a tool meant to be quick and trustworthy, that's a lot of friction.

The challenge: deliver real deep-learning classification without any patient data leaving the device, while staying fast and light enough to run comfortably in a normal web browser.

My role

Model, app, and deployment.

Modelling

Converted the trained model to ONNX, trading size against accuracy deliberately.

On-device inference

Wired up ONNX Runtime Web so the model executes fully client-side.

App

Built the React + Vite interface with upload, prediction, and real-time confidence scoring.

Deployment

Shipped it as a live, shareable app on Vercel.

Tech stack

What it's built with.

Model

YOLOv8ONNX

In-browser runtime

ONNX Runtime Web

Frontend

ReactVite

Hosting

Vercel
Architecture

Everything happens on your side of the wire.

The entire pipeline, from the moment a scan is selected to the confidence bars, executes inside the browser tab. There's no inference server to send data to.

IN THE BROWSER · NOTHING LEAVES THE DEVICE Upload MRIlocal file,in-memory Preprocessresize &normalise ONNX RT Web~1.5MB model~9ms inference 4-class + conf.softmax &confidence UI Server / cloud no upload
A fully client-side inference pipeline, private by design.
Key features

What it does.

  • Four-class classification of brain-MRI scans with a clear predicted type.
  • Fully in-browser: inference runs on-device through ONNX Runtime Web.
  • Privacy-first: scans are processed locally and never sent to a server.
  • Real-time confidence scoring so users see how sure the model is per class.
  • Featherweight & fast: a ~1.5 MB model returning results in about 9 ms.
  • Live & shareable: deployed on Vercel, no install required.
Challenges

The hard parts, and how I solved them.

On-device ML

Running a neural net inside a browser tab

Browsers aren't built for deep learning. I used ONNX Runtime Web with a model exported and trimmed to a ~1.5 MB footprint, which loads quickly and runs at roughly 9 ms per inference, fast enough to feel instant.

Privacy

Keeping data on the device

By moving inference to the client, the architecture removes the need to upload sensitive scans at all, privacy becomes a property of the design, not a policy bolted on afterwards.

Trade-offs

Small model, honest accuracy

Shrinking a model always costs something. I balanced size and speed against accuracy to land at a practical ~79%, a deliberate engineering trade-off for a tool that has to load and run anywhere, on any device, with no backend.

Results & impact

Fast, private, and live.

~79%
accuracy
~9ms
inference time
~1.5MB
model size
0
bytes of data uploaded

The finished app proves a full deep-learning workflow can live entirely on the client: private by construction, tiny, fast, and deployed for anyone to try.