About
Ten animals, one small model, running on your device.
Point it at a photo and it tells you which domesticated animal it sees. Nothing is uploaded โ the model itself is downloaded to your browser and does the thinking there.
Its entire vocabulary
- cat
- chicken
- cow
- dog
- duck
- goat
- horse
- pig
- rabbit
- sheep
Ten animals. It has never been shown anything else, so it will always answer with one of these โ or admit that it isn't sure.
How it learned
The model never learned to see from scratch. It starts from MobileNetV2, an existing model that had already studied more than a million everyday photographs and picked up the general grammar of images: edges, textures, fur, the shape of a leg.
That knowledge was frozen and left alone. A small new layer was added on top and trained only to tell these ten animals apart. Afterwards the last few layers of the original model were nudged, gently, to pay more attention to farmyards than to furniture.
This is called transfer learning. Borrowing a model that already knows how to look means you need a fraction of the photographs, and a fraction of the time, that starting from nothing would demand.
Where the photos came from
It learned from about 13,900 photographs โ roughly 1,200 to 1,450 of each animal โ gathered from Open Images, a Kaggle cats-and-dogs set, and web image search.
Raw photos are messy. Duplicates were removed, and every image was screened for things that were not what they claimed to be: cartoons, plush toys, watermarked stock photos, pictures that were mostly people. The most useful catch was a pile of sheep filed under goat. Left alone, it would have taught the model the wrong lesson twice over.
From training to the web
A freshly trained model is a Python object, and a browser cannot read one. So it was converted into a format the browser understands and packaged as a ~9 MB download that your browser fetches once, then keeps.
The converted model was then checked against the original on the same photographs, to confirm it still gives the same answers. A conversion that quietly changes the model's mind is exactly the kind of bug that never announces itself.
It runs on your device
There is no server doing the thinking. Once the model has downloaded, every prediction happens inside your browser tab, on your machine. Photos and webcam frames are never uploaded, never stored, and never leave your device.
What it gets wrong
It confuses cows and sheep more than anything else. At a distance, in a field, they share a silhouette โ and the training photos are honest about that.
When it is less than 60% sure, it says so instead of guessing confidently. And it does its best work on a clear, well-lit photo of a single animal; a dark, crowded, or half-hidden subject will throw it.