Skip to content

Densecap: dense video captions ​

Densecap is a continuous, fine-grained record of a first-person video. Examples show the desired specificity; they are not a closed vocabulary or action taxonomy.

Timeline and spans ​

  1. Ground every statement in the pixels of the presented view. Do not import facts from another view.
  2. Cover the video continuously in chronological order. Start at 00:00, end at the clip boundary, and leave no gaps or overlaps.
  3. Start a new span whenever the visible observation changes materially. Do not compress distinct movements, contacts, posture changes, or object effects into a coarse sentence; busy manipulation often needs 1–3 second spans.
  4. If the task supports 4 fps frame review with readable frame timestamps, use quarter-second steps .00, .25, .50, and .75. This is separate from VTG, which follows its prompt and rounds to the nearest second.
  5. Multiple sentences may appear under one span when every sentence describes a visible change. Do not repeat an unchanged observation to fill space.

Scene, objects, and space ​

  1. Establish the task-relevant scene state at the beginning and update layout, locations, and visible background objects when the camera or objects change meaningfully.
  2. Give exact counts only when every item is clearly visible and distinguishable. Use “several,” “at least,” or “partly obscured” when items overlap or are hidden.
  3. State object identity, color, material, and attributes only as supported by the image. Use “appears to be” or “looks like” when uncertain and keep references consistent.
  4. Use supported spatial language such as viewer’s left or right, clockwise or counterclockwise, along an edge, into a holder, or against a surface. Viewer-relative direction is the default; state any other reference frame.

Hands, posture, and mechanics ​

  1. Track left and right hands separately when anatomical identity is reliable. Otherwise describe the visible hand without guessing its side.
  2. Describe stepping, turning, bending, kneeling, rising, leaning, and paths toward or away from objects when posture or locomotion affects the action.
  3. Describe actual hand-object or hand-environment mechanics. Prefer concrete verbs over “arranges,” “handles,” or “manipulates.”
  4. Make the observable effect central: position, orientation, arrangement, shape, support, containment, or contact state.
  5. Combine causing action with before-and-after state, such as placing one package on another to form a stack of two.
  6. Describe approach, contact, maintained contact, support transfer, release, and resting contact when visible.
  7. Record meaningful pauses, hovering, bracing, and coordinated hand roles. Do not omit the stabilizing hand.

Uncertainty and repeated work ​

  1. Do not infer hidden contents, invisible forces, intent, speech, material properties, or successful fastening. State visible evidence and qualify uncertainty, for example: “The contact point is partly obscured, but the value changes from 28.0 to 28.5 while the gripper stays there, so it very likely presses the temperature control.” Do not write “the annotator believes” or “the annotator interprets”; put reasoning in a separate note if needed.
  2. Describe repeated cycles separately when their mechanics or outcomes are visible. Do not replace them with “continues,” and do not repeat unchanged observations.
  3. Include macro information inside chronological spans when visibly supported, such as a confidently countable number of cord loops. Otherwise say “several loops” or describe the coil thickening and loose cord shortening.
  4. Busy clips may use roughly 1,100–1,500 English words per minute as a reference, not a universal quota. Review factual grounding, counts, hand tracking, mechanics, effects, temporal granularity, and hallucinations.

Audio ​

Batch-0 video files do not include audio. Do not infer that the videos have audio from a mistaken example that mentions background speech.

Aboda VAS Help