Explainers··5 min read

How AI video search works, in plain English

What actually happens when you type “a city at night” and get back the right frames: sampling, embeddings, shots, transcripts and ranking.

1. Frames are sampled from the video

A video is thousands of still images. Indexing every single one would be slow and repetitive, because neighbouring frames are nearly identical. Instead, frames are sampled at regular intervals, which keeps enough detail to find any moment without wasting work on duplicates.

2. Each frame becomes an embedding

Each sampled frame goes through a vision model that outputs an embedding: a long list of numbers describing what the picture contains. Frames that look alike, such as two beach shots at sunset, end up with similar numbers. Frames that look different end up far apart.

FrameSeek uses Azure AI Vision for this step, and stores the embeddings in a vector index so they can be compared quickly.

3. Your words land in the same space

When you search, your text goes through a matching model that places it in the same space as the images. “A city at night” lands near frames of lit-up streets and skylines, whatever the files are called and whether or not anyone tagged them.

Search then ranks frames by how close they are to your query and returns the best matches first. That is why there is no special syntax to learn: you describe the picture the way you’d describe it to a friend.

4. Similar frames are grouped into shots

Ten near-identical frames from the same moment aren’t ten results. FrameSeek groups consecutive frames that look alike into shots, and groups search results by video and shot, so you see distinct moments rather than a wall of duplicates. Shots also make clipping easier, because a shot is usually the clip you want.

5. Audio is transcribed separately

Image embeddings capture what is on screen, not what is said. To cover speech, the audio track is transcribed (FrameSeek uses Azure OpenAI for this). Transcript lines carry timestamps, so finding a phrase in the transcript takes you to the right moment in the video.

What this means for your searches

  • Describe what you would see, not what happened before or after.
  • Concrete nouns, settings, colours and actions match best.
  • Use the transcript for dialogue, names and anything that was spoken.
  • Nothing needs to be tagged by hand. A video is searchable as soon as processing finishes.

Written by the FrameSeek team.

A little less searching. A little more creating.

Find the moment. Make the thing.

Sign in with Google and search your first video in minutes. Free plan, no card.