Computer vision · 7 min read
What an image embedding is and what it lets you find
A frame becomes a list of numbers where similar scenes sit close together. That finds the duplicates, the rare scene, and the night shift missing from your set.
Summary
This post explains an image embedding as a list of numbers in which similar frames sit close together, and shows what that lets a team find in an archive: near-duplicates, a rare scene described in words, and the conditions missing from a training set. It concludes that an embedding finds frames and a trained detector finds things in them, and that the two jobs stay separate. It is for anyone with a large archive of footage and a small labeling budget.
Sheikh Srijon · GTM Lead · Sep 28, 2026

Thermal view of a yard at night, seven people boxed among bare trees, from a customer site camera
A yard camera on a distribution site has recorded for eleven months, and a team wants to train a model on it. The archive is enormous and almost entirely the same frame: an empty yard, a fence, a bare tree, at every hour of every day. Somewhere in it are the few hundred frames that matter, the night the fence was climbed, the week in January the snow came, the afternoon a contractor's van parked where nothing parks. Nobody can watch eleven months of footage to find them.
An embedding is the tool for that job. It turns each frame into a list of numbers such that frames that look alike, in a sense the model learned, end up close together, and frames that differ end up far apart. Once every frame in the archive is a point in that space, questions that were impossible become cheap.
A frame becomes a point where similar scenes sit close
An embedding model takes a frame and returns a few hundred numbers. On their own the numbers mean nothing. What means something is distance: two frames of the empty yard at 3 pm on consecutive days produce nearly identical lists, and a frame of the yard with a van in it produces a list that sits some way off. The space is arranged by content rather than by pixels. A frame of the yard at night and a frame of it at noon share few pixel values, and still sit closer to each other than either does to a frame of a warehouse aisle.
Distance is usually measured as the cosine of the angle between two lists, which ignores their length and keeps their direction. A score near one is near-identical. The threshold that counts as "the same scene" is chosen by looking at pairs, and it differs from one camera to the next.
Contrastive learning is how the space gets its shape
The model learned the arrangement by being shown pairs. Contrastive learning trains a model to pull matching pairs together and push mismatched ones apart. The best-known versions, CLIP among them, pair each image with a caption, so that the frame and the words "a van in a yard at night" land near each other. That is why a description in plain words can be used to search: the words become a point in the same space, and the frames nearest that point are the answer.
The pairs came from the internet, which is where the model's sense of "alike" comes from. It has a strong opinion about vans and fences and snow, and a weak one about the difference between a worn and a new fence panel, because nobody captioned that. The space is a good map of what people write about pictures and a poor map of what an inspector looks for.
Near-duplicates are the first thing an embedding finds
The most immediate use is subtraction. Eleven months at one frame every couple of seconds is an enormous number of frames, and nearly all of them are copies of their neighbours. Cluster the points, keep one frame per cluster, and the archive collapses to the few thousand frames that are actually different from each other. That is the set worth labeling, and the lesson on what a vision model can and cannot see makes the point from the other side: a model learns from distinct conditions, and near-duplicates teach it one condition many times over.
The same distance check runs on the way in. A frame from Tuesday that sits on top of one from Monday adds nothing, and can be dropped before anyone spends a second on it.
The rare scene can be found by describing it
Because words and frames share the space, the team can type "van parked by the gate at night" and get back the frames nearest that phrase, ranked. The contractor's van from March turns up in the first page. So does a night in July when a forklift was left by the gate, which nobody remembered and which is exactly the kind of frame the training set was missing.
This is search, and it is a different thing from detection. The search returns frames that resemble the description; it draws no box, and it is wrong in ways a detector is not. A frame of a white sign by the gate can rank above a dark van. The person reading the results is part of the method, and a rare scene found this way still has to be labeled before it teaches anything.
The Sunday frames on that camera, incidentally, cluster on their own, because the site's one Sunday delivery arrives in a truck of a colour nothing else on the site has.
The empty region of the space is the night shift you never labeled
Turn the map around and look for where the labeled frames are absent. Plot the training set's points over the archive's points, and the archive has regions with no training frame anywhere near: the snow week, the fog mornings in November, the night frames from the thermal camera. Those gaps are the conditions the model has never seen, and they are visible before the model is trained rather than after it fails.
It also works on the live feed. A frame from the camera that lands far from everything in the training set is a frame the training set does not cover, and that is a reasonable thing to send to a person even before a detector has had its say. It is one signal among several, and the signal that decides a model has drifted is still the rate at which people correct it.
An embedding finds frames and a detector finds things in them
The two jobs stay separate, and the separation is the honest boundary of the tool. An embedding can say "this frame is like that one" and "this frame is like these words". It cannot say where in the frame the van is, how many people are at the fence, or whether the panel that looks like the others is the one with the crack. A trained detector does those, on frames a person has labeled, and the embedding's job is to choose which frames get labeled.
LexData takes that detector through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the site already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The labeling doc covers how a prompt becomes a class list; the embedding is what decides that the March van and the July forklift are in the set the prompt runs on.
My own view is that the embedding earns more of its keep before training than after. Finding the eleven-month archive's few thousand distinct frames, and the night shift missing from them, is where a small labeling budget goes furthest, and a team that skips it labels the same empty yard a thousand times.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
When the outline is the answer and a box will not do
A weld defect whose area decides the rework and a part a robot has to grasp both need a mask. Here is what the mask costs to label and when a box is enough.
Esdras Ntuyenabo · Sep 28, 2026

Computer vision · 6 min read
DETR explained, what the detection transformer changed about finding objects
Detection as set prediction, a fixed set of queries instead of anchors, no suppression step, and the small-object and slow-training problems later fixed.
Finn Ellingwood · Sep 28, 2026

Computer vision · 6 min read
Reading a training run before trusting the model
The loss curve says the run finished. The gap to validation, the recall on the crack class and the frames behind the score say whether the line can trust it.
Rob Hickey · Sep 28, 2026