Your tech walks the site with the camera running and says what they see. Fieldproxy reads the footage and the narration together, and writes the survey, the inspection, or the damage report.
Photos tell you what was there. They cannot tell you the gate code, the dog in the back garden, or the thing your tech muttered about the flashing on the way past. The narration carries all of that, and it is read as part of the same clip.
Four steps, and the fourth is the one that matters.
Step 1
Camera running, narrating as they go. On mobile they can record it right inside the job; on web they pick a clip that already exists.
Step 2
Up to 12 frames, evenly spaced, each labelled with its timestamp so the model reads one continuous walkthrough rather than a pile of unrelated photos.
Step 3
What your tech said out loud is the half a photo can never capture: the gate code, the dog, the thing that is wrong but not visible.
Step 4
Not frames then audio. Shown and said are read as one thing, then written into your fields, your dropdowns, your grid rows.
The frames are sampled and the audio is extracted on the phone or in the browser itself. Only a handful of stills and a small audio file are sent for reading. The footage stays where it was recorded.
That matters for site footage more than most things: a walkthrough of a customer's premises records their property, their staff and whatever is lying around. If you do want the clip kept, because the recording is the evidence, you turn that on deliberately and one pick both reads the video and stores it against the job.
No asterisks. These are the real limits.
Up to 12 frames
Sampled evenly across the clip. Six by default.
Clips up to 5 minutes
Longer than that and it asks for a shorter one rather than guessing.
Narration read to 3 minutes
A longer clip is still sampled for frames across its full length; only the transcript is cut, and it tells you so.
Silence is fine
No audio track, or one that cannot be decoded, and it carries on from the pictures alone and says which.
Roof type, storeys, access notes, what is visibly wrong. The walkthrough replaces a clipboard and a second visit.
Turn on store_video and the timestamped clip is kept alongside the filled form. For dispute work the recording IS the deliverable: it is what you show the customer to justify the charge.
One pass of the finished job, narrated. Serial numbers, positions, exceptions, and the commissioning checklist filled from the same clip.
In rows mode the clip fills an editable grid, a row per item, rather than a single set of fields.
AI and humans, same job
The clip fills the form. Your tech still checks it before it is saved, and every field says how confident it was. Neither is doing the other's work.
Most teams adopt one agent first, then layer in two or three more in the same week.
Works in every trade