Skip to content

FAQ ​

This page collects common questions from regular annotators and the Batch-0 QA record.

System operations ​

Why can’t I see a task? ​

Check that you passed learning and the exam for that type and batch. Qualifications are independent; passing EQA does not unlock VTG. Check the batch and type filters, then contact the administrator.

What happens after an incorrect practice answer? ​

It does not count toward learning progress. Use Practice again after reviewing the result. Repeating a correct question does not add another count.

Why can’t I start an exam? ​

Learning may be incomplete, the required number of questions may not be selected, the unpractised pool may be too small, or the attempt limit may be reached. Follow the page message and contact the administrator when the pool is insufficient.

What if saving fails or a video will not play? ​

Check the network and refresh. Record the batch, task ID, and error message. Do not guess and do not edit the same task in multiple tabs.

Can I edit after submission? ​

Wait for review. If returned, read the reason in Logs or Versions, revise, save, and submit again.

EQA questions ​

Are formatting, omissions, and action errors scored separately? ​

When a task specifies a format, follow it exactly. The answer must satisfy the prompt, its format, and its visible evidence. Do not add unrequested explanations.

What if the video is clear but the EQA answer is still uncertain? ​

Some prompts allow more than one reasonable interpretation. Choose the answer best supported by the video, stay consistent within the item, and explain the rationale in a separate note when available. Do not invent unseen information.

The prompt says robot but the video shows a person. Should I edit the prompt? ​

Do not edit it. Flag the prompt-video mismatch for review.

The prompt says zucchini but the visible objects look like cucumbers. Should I answer zero? ​

There are no trick questions. Answer the visible target count, such as 7 in that example, and flag that the prompt should say cucumber.

The prompt asks for rice bags but the four visible bags are tissue. What should I do? ​

Answer 4 when the event is that four bags were put into the cart, and flag that the prompt should ask about tissue bags. Do not treat the language on the packaging as a trick condition.

How should I interpret an ambiguous “fold” prompt? ​

Use the action specified by the surrounding prompt and the visible sequence. If the question refers to the fold after the dumpling is enclosed, that interpretation is reasonable; record the ambiguity in a note.

Should the first stair before the stated range be counted? ​

Follow the wording. If the prompt does not include the first step before the stairs, leaving it out is correct.

VTG questions ​

Should 1.4 seconds be 1 or 2 seconds? ​

Round to the nearest second as required by VTG: 1.4 seconds is 00:01.

Does a brief loss of contact followed by immediate contact count as a new event? ​

Use the event definition in the prompt. When there is no minimum separation rule, apply your own consistent interpretation and explain the rationale in a note if possible.

The same cap is picked up twice. How many timestamps? ​

Record both when the prompt asks for every pick-up. Do not merge repeated events.

A screw is already on the screwdriver and only moves to the tip. Is that a first attachment? ​

Count only the first contact defined by the prompt. If “first attached” is mixed with “moved to the tip,” flag the ambiguity. A clearer prompt would ask for the moment the screw first touches any part of the screwdriver.

A hand turns twice quickly. Is the second movement another turn or just stabilizing? ​

If the hand leaves the turntable and a second rotation is visible, record “turning again” as a new anchor. Use the visible action and prompt definition.

What if timestamps seem slightly off? ​

Identify the event itself, such as the moment a stick enters a die, rather than the hand approaching or leaving. Then apply nearest-second rounding. Flag the task when the event definition is unclear.

Densecap questions ​

How should I describe an unidentifiable black object on the floor? ​

Describe only visible evidence and useful spatial relationships, such as “a black object on the floor in front of the round pedestal table.” Do not suggest that it may be a backpack. When visible, relationships such as “a lined trash bin just outside the doorway” are useful.

Can I say that a gripper pressed a control when the contact point is hidden? ​

Do not write “the annotator believes” or “the annotator interprets.” State the visible evidence and qualify uncertainty: “The gripper remains against the lower-right control area; the exact contact point is partly obscured, but it very likely presses the temperature control because the display changes from 28.0 to 28.5 while the gripper remains there.” Put reasoning in a separate note when needed.

Can one timestamp span contain multiple sentences? ​

Yes, when every sentence describes a visible change in that span. Do not repeat an unchanged observation.

Should Densecap include macro information such as cord-loop counts? ​

Include an exact count only when it is clearly visible and confidently countable. Otherwise say “several loops” or describe the coil becoming thicker and the loose cord shortening. Keep macro information inside chronological spans and retain the distinguishable manipulation phases.

What level of detail is expected? ​

A strong span can state both hand roles and the changing object state: “From 00:10 to 00:16, the left hand steadies and rotates the dryer while the right hand winds the remaining cord around the handle, leaving the plug end loose; the right hand then gathers the plug end and begins tucking it under the coils.” One span is acceptable when a transition is gradual or cannot be timed reliably, provided both phases remain explicit.

What timestamp precision should Densecap use? ​

If the video can be reviewed at 4 fps with visible frame timestamps, use .00, .25, .50, and .75 steps. This does not replace the VTG rule: VTG timestamps are rounded to the nearest second.

Do the source videos have audio? ​

No. Batch-0 video files do not include audio. A gold example that mentions background speech contains an annotation mistake.

Which direction convention should I use? ​

Use the viewer’s perspective by default for left, right, and turning direction. State another reference frame explicitly.

Contact ​

WeChat support: zerk_c

Include your username (never your password), batch and task ID, current page, steps taken, and an error message or screenshot with sensitive information masked.

Aboda VAS Help