Video Agent

One walkthrough video. The whole form filled.

Your tech walks the site with the camera running and says what they see. Fieldproxy reads the footage and the narration together, and writes the survey, the inspection, or the damage report.

Photos tell you what was there. They cannot tell you the gate code, the dog in the back garden, or the thing your tech muttered about the flashing on the way past. The narration carries all of that, and it is read as part of the same clip.

How the Video Agent works

Four steps, and the fourth is the one that matters.

Step 1

Your tech walks and talks

Camera running, narrating as they go. On mobile they can record it right inside the job; on web they pick a clip that already exists.

Step 2

Stills are sampled across the clip

Up to 12 frames, evenly spaced, each labelled with its timestamp so the model reads one continuous walkthrough rather than a pile of unrelated photos.

Step 3

The narration is transcribed

What your tech said out loud is the half a photo can never capture: the gate code, the dog, the thing that is wrong but not visible.

Step 4

Both go to the model together

Not frames then audio. Shown and said are read as one thing, then written into your fields, your dropdowns, your grid rows.

The video never leaves the device

The frames are sampled and the audio is extracted on the phone or in the browser itself. Only a handful of stills and a small audio file are sent for reading. The footage stays where it was recorded.

That matters for site footage more than most things: a walkthrough of a customer's premises records their property, their staff and whatever is lying around. If you do want the clip kept, because the recording is the evidence, you turn that on deliberately and one pick both reads the video and stores it against the job.

What it does, exactly

No asterisks. These are the real limits.

Up to 12 frames

Sampled evenly across the clip. Six by default.

Clips up to 5 minutes

Longer than that and it asks for a shorter one rather than guessing.

Narration read to 3 minutes

A longer clip is still sampled for frames across its full length; only the transcript is cut, and it tells you so.

Silence is fine

No audio track, or one that cannot be decoded, and it carries on from the pictures alone and says which.

Where teams point it

Site surveys

Roof type, storeys, access notes, what is visibly wrong. The walkthrough replaces a clipboard and a second visit.

Damage and disputes

Turn on store_video and the timestamped clip is kept alongside the filled form. For dispute work the recording IS the deliverable: it is what you show the customer to justify the charge.

Install and handover

One pass of the finished job, narrated. Serial numbers, positions, exceptions, and the commissioning checklist filled from the same clip.

Multi-line inspections

In rows mode the clip fills an editable grid, a row per item, rather than a single set of fields.

AI and humans, same job

The clip fills the form. Your tech still checks it before it is saved, and every field says how confident it was. Neither is doing the other's work.

Bring the form you fill after every site visit

On the mapping call we walk your process end to end and find where this fits. Then we configure it on your workflows and you watch it run.

Part of the Fieldproxy platform

Every agent runs on the same AI Agents stack

Voice, vision, dispatch, and knowledge agents all share one runtime, one schema, and one customization surface, describe what you want in plain English.

See it running on your operation

Book a 20-minute demo and we'll show you this agent (and the rest of our AI Agents stack) tailored to your real workflows.