Dataset
Dataset
RefCOCO Dataset
In the ReferenceCOCO dataset, an image containing segmented target objects is paired with its corresponding natural language expression referring to the target object. Data collected for this dataset is intended to produce a referring expression dataset that focuses purely on appearance-based description, e.g., "the man in the yellow polka-dotted shirt" rather than "the second man from the left".
RefCOCO consists of 142,209 refer expressions for 50,000 objects in 19,994 images.
RefCOCO+ has 141,564 expressions for 49,856 objects in 19,992 images.