computer vision dataset

Search Computer vision data

For fast search follow this website http://www.computervisiondatasets.ml/

A curated list of awesome computer vision datasets.

Contributing and Collaborating

Please feel free to send me pull requests or email ([email protected]) to add links.

Object recognition (also called object classification)

The classes are completely mutually exclusive. There is no overlap between automobiles and trucks.

Image Classification

MNIST

Classes: 10 (0-9)
Training: 60,000
Test : 10,000
Image type: handwritten digits
It is a subset of a larger set available from NIST.

CIFAR-10 and CIFAR-100

CIFAR-10

Classes: 10 (airplan,automobile,bird,cat,deer,dog,frog,horse,ship,truck)
Training: 50000 (5000 images per class)
Test: 10000 (1000 per class)
Image type: airplan,automobile,bird ...
It is a subset of a larger set available from 80 million tiny images.

CIFAR-100

Classes: 100 (airplan,automobile,bird,cat,deer,dog,frog,horse,ship,truck...)
Training: 50000 (500 images per class)
Test: 10000 (100 per class)
Image type: airplan,automobile,bird ...
It is a subset of a larger set available from 80 million tiny images.

Note: The classes are completely mutually exclusive. There is no overlap between automobiles and trucks. "Automobile" includes sedans, SUVs, things of that sort. "Truck" includes only big trucks. Neither includes pickup trucks.

STL-10

Classes : 10 (airplane, bird, car, cat, deer, dog, horse, monkey, ship, truck)
Training : 500 training images (10 pre-defined folds)
Test : 800 test images per class
Image : airplane, bird, car ...
100000 unlabeled images for unsupervised learning.
Images were acquired from labeled examples on ImageNet.

Note : It is inspired by the CIFAR-10 dataset but with some modifications. In particular, each class has fewer labeled training examples than in CIFAR-10, but a very large set of unlabeled examples is provided to learn image models prior to supervised training. The primary challenge is to make use of the unlabeled data (which comes from a similar but different distribution from the labeled data) to build a useful prior.

The Street View House Numbers (SVHN) Dataset

Classes : 10 classes, 1 for each digit. Digit '1' has label 1, '9' has label 9 and '0' has label 10
Training : 73257 digits for training
Testing : 26032 digits for testing,
Image type : hous number digits
Additional : 531131 additional, somewhat less difficult samples, to use as extra training data

Note : It can be seen as similar in flavor to MNIST (e.g., the images are of small cropped digits), but incorporates an order of magnitude more labeled data (over 600,000 digit images) and comes from a significantly harder, unsolved, real world problem (recognizing digits and numbers in natural scene images). SVHN is obtained from house numbers in Google Street View images.

ILSVRC2012 task 1

Classes : 1000
With tens of thousands of training, validation and testing images.
Image type : airplane, bird, car ....
he training data, the subset of ImageNet containing the 1000 categories and 1.2 million images,

Note : The 1000 object categories contain both internal nodes and leaf nodes of ImageNet, but do not overlap with each other.

PASCAL VOC 2009 dataset

Classes : 20 ( persons, animals(bird, cat, cow, dog, horse, sheep),Vehicle(aeroplane, bicycle, boat, bus, car, motorbike, train),Indoor(bottle, chair, dining table, potted plant, sofa, tv/monitor))
Image type : airplane, bird, car...

Note : Pascal contain Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets

Caltech 101

Classes : 101 categories (About 40 to 800 images per category. Most categories have about 50 images.)(e.g., “helicopter”, “elephant” and “chair” etc.)
totalling around 9k images.
Image type : “helicopter”, “elephant” and “chair” etc.

Caltech 256

Classes : Pictures of objects belonging to 256 categories
Total 30,607 real-world images,
Image type : “helicopter”, “elephant” and “chair” etc.

Flower classification data sets

17 Flower Category Dataset

Classes : 17 (80 images for each class)(common flowers in the UK.)
Image type : flowers

102 category dataset

Classes : 102 (Each class consists of between 40 and 258 images.)(common flowers in the UK.)
Image type : flowers

Animals with attributes 2

Classes : 50 (animal classes)
Total 30475 images
Image type : animals

Stanford Dogs Dataset

Classes : 120
Number of images: 20,580
Image type : dogs

McGill Real-World Face Video Database

Classes : undefined (this dataset particularly focus on head pose ground truth and estimate facial attributes but we can also use in image classification)
This database contains 18000 video frames of 640x480 resolution from 60 video sequences, each of which recorded from a different subject (31 female and 29 male).
Image type : face / face expression

mage Classification: People & Food

Classes: unknown (people eating fruits, cakes, and other foodstuffs.)
Image types : people eating fruits, cakes, and other foodstuffs.

Images of Crack in Concrete for Classification

Classes : 2 (negative and positive crack images)
Total data points : 40,000 images of concrete
Image type : Concrete Crack Images

Architectural Heritage Elements

Classes : 10 categories: Altar: 829 images; Apse: 514 images; Bell tower: 1059 images; Column: 1919 images; Dome (inner): 616 images; Dome (outer): 1177 images; Flying buttress: 407 images; Gargoyle (and Chimera): 1571 images; Stained glass: 1033 images; Vault: 1110 images.
Total datapoints : 10235 images
Image type : Architectural Heritage Elements

Fruits 360

classes: 131 (fruits and vegetables).
Training set size: 67692 images (one fruit or vegetable per image).
Test set size: 22688 images (one fruit or vegetable per image).
The total number of images: 90483.
Image type : fruit or vegetable

Medical Images

TensorFlow patch_camelyon Medical Images

Classes : 2(binary)
Containing over 327,000 color images
Image type : Histopathologic Cancer Detection. Identify metastatic tissue in histopathologic scans of lymph node sections.

Recursion Cellular Image Classification

Classes : Your entry will classify images of cells under one of 1,108 different genetic perturbations.
Image type: Cellular Images

Blood cell images

Classes : 4 (The cell types are Eosinophil, Lymphocyte, Monocyte, and Neutrophil.)
Total images : 12,500 augmented images(3,000 images for each of 4 different cell types)
Image type : The cell types are Eosinophil, Lymphocyte, Monocyte, and Neutrophil.

ChestX-ray8

Classes : 8 (eight common disease labels)
Total : 108,948 frontal-view X-ray images of 32,717 unique patients collected between 1992 and 2015.
Image type : Chest X-ray

Agriculture and Scene

Indoor Scenes Images

Classes : 67 separate categories(kitchen,operating_room,restaurant_kitchen...)
Total : 15,000+ images
Image type : kitchen,operating_room,restaurant_kitchen...

Images for Weather Recognition

Classes : 4 separate categories based on sunrise, cloudy, rainy, and sunshine.
Total images :
Image types : sunrise, cloudy, rainy, and sunshine

Intel Image Classification

Classes : 6 categories(buildings,forest,glacier,mountain,sea,street)
Total datapoints : 25,000
Type of image : buildings,forest,glacier,mountain,sea,street

TensorFlow Sun397 Image Classification Dataset

Classes : 397 (house,outdore,station,playground...)
Total datapoints : 108,000 (The number of images varies across categories, but there are at least 100 images per category.)
Type of image : house,outdore,station,playground

Video Classification

Video classification and image classification models both use images as inputs to predict the probabilities of those images belonging to predefined classes. However, a video classification model also processes the spatio-temporal relationships between adjacent frames to recognize the actions in a video.
For each frame, pass the frame through the CNN, Classify each frame individually and independently of each other
Action dataset are mostly use in video classification and video classification data can be use in image classification

Video classification USAA dataset

Classes : 8
Total videos : 100 (Each video is labeled by 69 attributes and 69 attributes can be broken down into five broad classes: actions, objects, scenes,sounds, and camera movement.)
Type of videos : home videos of social occassions which feature activities of group of people

UCF50

Classes : 50
Type of videos: Baseball Pitch, Basketball Shooting, Bench Press, Biking, Biking, Billiards Shot,Breaststroke, Clean and Jerk, Diving, Drumming,

Note : Extension of YouTube Action data set

UCF 101

Classes(Actions) : 101
Type of videos : Baseball Pitch, Basketball Shooting, Bench Press, Biking, Biking, Billiards Shot,Breaststroke, Clean and Jerk, Diving, Drumming,

Note : Extension of UCF50

Automobile and ADAS Related Datasets

German Traffic Sign Classification Dataset

References

Detection(object detection,Edge detection)/ Multi-view Object Detection / Segmentation(Image Segmentation) / Saliency(salient) Detection / Semantic labeling

e-Lab Video Data Set

Classes : 35 (plant,shoes... objects in our environment)
Total videos : 2050 videos of roughly 10 seconds each.

Detection

COCO

330K images (>200K labeled),1.5 million object instances,80 object categories,91 stuff categories,5 captions per image

KITTI

The object detection and object orientation estimation benchmark consists of 7481 training images and 7518 test images, comprising a total of 80.256 labeled objects.

ImageNet

14,197,122 images and 1,000 categories

BDD100K

largest open driving video dataset as part of the CVPR

DOTA

Dataset for Object deTection in Aerial Images , 11268 Items ,18 Classes ,

MaskedFaceNet

MaskedFace-Net is a dataset of human faces with a correctly or incorrectly worn mask (133,783 images) based on the dataset Flickr-Faces-HQ (FFHQ).

CIFAR-10 -The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.

LISA Traffic Sign Detection Dataset

The LISA Traffic Sign Dataset is a set of videos and annotated frames containing US traffic signs.

Exclusively Dark (ExDark) Image Dataset

12 class , 7363 Labels

The 20BN-SOMETHING-SOMETHING Dataset V2

Total number of videos 220,847/ Training Set 168,913,Validation Set 24,777 / Test/ Set (w/o labels) 27,157 / Labels 174 / Quality 100px / FPS 12

PASCAL VOC 2007,2009,2010,2011 dataset

Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets

LabelMe dataset

LabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you use the database, we only ask that you contribute to it, from time to time, by using the labeling tool.

CMU/VASC & PIE Face dataset

The CMU Multi-PIE face database contains more than 750,000 images of 337 people recorded

Yale Face dataset

The Yale Face Database (size 6.4MB) contains 165 grayscale images in GIF format of 15 individuals.

Caltech

Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds

Caltech 101

Pictures of objects belonging to 101 categories

Caltech 256

Pictures of objects belonging to 256 categories

Daimler Pedestrian Detection Benchmark

15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set contains more than 21,790 images with 56,492 pedestrian labels (fully visible or partially occluded), captured from a vehicle in urban traffic.

MIT Pedestrian dataset

CVC Pedestrian Datasets

CVC Pedestrian Datasets

CBCL Pedestrian Database

MIT Face dataset

CBCL Face Database

MIT Street dataset

CBCL Street Database

INRIA Person Data Set

A large set of marked up images of standing or walking people

INRIA car dataset

A set of car and non-car images taken in a parking lot nearby INRIA

INRIA horse dataset

A set of horse and non-horse images

H3D Dataset

3D skeletons and segmented regions for 1000 people in images

HRI RoadTraffic dataset

A large-scale vehicle detection dataset

BelgaLogos

10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.

FlickrBelgaLogos

10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.

FlickrLogos-32

The dataset FlickrLogos-32 contains photos depicting logos and is meant for the evaluation of multi-class logo detection/recognition as well as logo retrieval methods on real-world images. It consists of 8240 images downloaded from Flickr.

TME Motorway Dataset

30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated using laser-scanner data. Distance estimation and consistent target ID over time available.

PHOS (Color Image Database for illumination invariant feature selection)

Phos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 15 different images: 9 images captured under various strengths of uniform illumination, and 6 images under different degrees of non-uniform illumination. The images contain objects of different shape, color and texture and can be used for illumination invariant feature detection and selection.

CaliforniaND: An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections

California-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate cases, without the use of artificial image transformations. The dataset is annotated by 10 different subjects, including the photographer, regarding near duplicates.

USPTO Algorithm Challenge, Detecting Figures and Part Labels in Patents

Contains drawing pages from US patents with manually labeled figure and part labels.

Abnormal Objects Dataset

Contains 6 object categories similar to object categories in Pascal VOC that are suitable for studying the abnormalities stemming from objects.

Human detection and tracking using RGB-D camera

Collected in a clothing store. Captured with Kinect (640*480, about 30fps)

Multi-Task Facial Landmark (MTFL) dataset

This dataset contains 12,995 face images collected from the Internet. The images are annotated with (1) five facial landmarks, (2) attributes of gender, smiling, wearing glasses, and head pose.

WIDER FACE: A Face Detection Benchmark

WIDER FACE dataset is a face detection benchmark dataset with images selected from the publicly available WIDER dataset. It contains 32,203 images and 393,703 face annotations.

PIROPO Database: People in Indoor ROoms with Perspective and Omnidirectional cameras

Multiple sequences recorded in two different indoor rooms, using both omnidirectional and perspective cameras, containing people in a variety of situations (people walking, standing, and sitting). Both annotated and non-annotated sequences are provided, where ground truth is point-based. In total, more than 100,000 annotated frames are available.

The Boxy vehicle detection dataset

A vehicle detection dataset with 1.99 million annotated vehicles in 200,000 images. It contains AABB and keypoint labels.

The Bosch Small Traffic Lights Dataset

A dataset for traffic light detection, tracking, and classification.

DriveU Traffic Light Dataset (DTLD)

It contains more than 40.000 images and 230 000 annotated traffic lights and is the largest database for traffic light detection so far containing bounding box labels, track identities and furthermore the following attributes: phase, pictogram, relevancy, occlusion, number of light units and orientation.

ETH (ETH Pedestrian)

ETH is a dataset for pedestrian detection. The testing set contains 1,804 images in three video clips. The dataset is captured from a stereo rig mounted on car, with a resolution of 640 x 480 (bayered), and a framerate of 13--14 FPS.

TUD-Brussels Pedestrian

In this experiment, we only use 288 images which contain pedestrians . TUD-Brussels test set contains 508 images containing 1498 annotated pedestrians.

The 2D Shape Structure Dataset

The 2D Shape Structure database is a public, user-generated dataset of 2D shape decompositions into a hierarchy of shape parts with geometric relationships reta...

ICS-FORTH + Modelling of 2D Shapes with Ellipses

The dataset contains more than 4,536 2D shapes included in standard as well as in home-build datasets. Our goal is to represent a given 2D shape with an au...

Mobile Phone and Webcam Hand Images for Personal Authentication and Identification

This work attempts to provide two Hand Images Databases for hand biometrics: one is created using a mobile phone camera of modest quality, which we called mob...

Detail 2D Projection DataSet

Detail 2D Projection DataSet is a database of 2d projections of mechanical details with holes. The dataset consists of 13 shape categories where each category i...

3DVis

The 3DVis dataset includes a set of 12 heterogeneous scenes for testing 3D scene registration and analysis methods. Models include homogeneous shapes, repetitiv...

PASCAL Context

We would like to announce the release of PASCAL-Context dataset. We augmented PASCAL VOC 2010 dataset with annotations for 400+ additional categories. In the cu...

SHOT 3D shape description

The 3D shape description dataset consists of multiple sub-datasets Descriptor Matching - Dataset 1 & 2 (Stanford) These datasets, created from some of the m...

THUR15000

We introduce a labeled dataset of categorized images for evaluating sketch based image retrieval. Using Flickr, we downloaded about 3000 images for each of the ...

RGB-D Person Re-identification

The RGB-D Person Re-identification dataset is for person re-identification using depth information. The main motivation is that the standard techniques (such as...

TUD Shapes 1+2

This material is supplementary to Michael Stark, Bernt Schiele. How Good are Local Features for Classes of Geometric Objects. Eleventh IEEE International C...

EITZ Sketch Quality

Humans have used sketching to depict our visual world since prehistoric times. Even today, sketching is possibly the only rendering technique readily available ...

EITZ Sketch-Based Image Retrieval

We introduce a benchmark for evaluating the performance of large scale sketch-based image retrieval systems. The necessary data is acquired in a controlled user...

ICG Sketch Retrieval

The ICG Sketch Retrieval dataset consists of XXX hand-drawn sketches for five categories. It is used for content-based image retrieval using shape features for ...

Simpsons 40 years

Simpsons Homer 40 years is a dataset showing Homer Simpson over the course of 40 years. It is used for video segmentation and shape matching between frames....

Leaves -The Leaves dataset from X contains X images of leaves. Leaves dataset taken by Markus Weber. California Institute of Technology PhD student under Pietro Per...

ETHZ Shape

The ETHZ Shape classes dataset from Vittorio Ferrari [?] consists of five object classes and a total of 255 images. All classes contain significant intra-class ...

Weizmann Horses

The multi-scale Weizmann horses (originally from Eran Borenstein, adapted by Jamie Shotton) consists of 656 images which is split into 50+50training, 50+50 vali...

ETHZ Extended Shape

The ETHZ Extended Shape classes dataset from Konrad Schindler is larger dataset of shape categories, created by merging ETHZ shape classes with Konrad Schindler...

Tools2D

The Tools 2D dataset from Bronstein, Bronstein, Bruckstein, and Kimmel [?] for partial similarity experiments and consists of 15 shapes: 5 humans, 5 horses and ...

Mythological Creatures -The Mythological Creatures consists of articulated shapes (silhouettes) for partial similarity experiments and contains 15 shapes: 5 humans, 5 horses and 5 cent...

SIID

The SIID silhouette dataset contains... and is from the Shape Indexing of Image Database (SIID). Download SIID silhouette dataset http://www.lems.brown.edu/...

KIMA216

The Kimia 216 has 18 classes each consisting of 12 images. It contains shapes silhouettes for birds, bones, brick, camels, car, children, classic cards, elephan...

KIMA99

The Kimia 99 has 9 classes each consisting of each 11 images. They are part of the Shape Indexing of Image Database (SIID) project, which also contains the SIID...

KIMIA25

The Kimia 25 consists of 6 classes and 25 images. They are part of the Shape Indexing of Image Database (SIID) project, which also contains the SIID silhouette ...

MPEG-7 Core Experiment CE-Shape-1

MPEG-7 Core Experiment CE-Shape-1 [?] is a popular database for shape matching evaluation consisting of 70 shape categories, where each category is represented ...

Labeled and Annotated Sequences for Integral Evaluation of SegmenTation Algorithms

ChangeDetection.net

Multi-view Object Detection

Saliency Detection

Salient Object Detection benchmark

This benchmark aggregates results from 36 methods over five datasets (MSRA10K, ECSSD, THUR15K, JuddDB, and DUTOMRON).

AIM

120 Images / 20 Observers (Neil D. B. Bruce and John K. Tsotsos 2005).

LeMeur

27 Images / 40 Observers (O. Le Meur, P. Le Callet, D. Barba and D. Thoreau 2006).

Kootstra

100 Images / 31 Observers (Kootstra, G., Nederveen, A. and de Boer, B. 2008).

DOVES

101 Images / 29 Observers (van der Linde, I., Rajashekar, U., Bovik, A.C., Cormack, L.K. 2009).

Ehinger

912 Images / 14 Observers (Krista A. Ehinger, Barbara Hidalgo-Sotelo, Antonio Torralba and Aude Oliva 2009).

NUSEF

758 Images / 75 Observers (R. Subramanian, H. Katti, N. Sebe1, M. Kankanhalli and T-S. Chua 2010).

JianLi 235 Images / 19 Observers (Jian Li, Martin D. Levine, Xiangjing An and Hangen He 2011).

Extended Complex Scene Saliency Dataset (ECSSD)

ECSSD contains 1000 natural images with complex foreground or background. For each image, the ground truth mask of salient object(s) is provided.

Pedestrian Detection

Caltech Pedestrian Detection Benchmark

ETHZ Pedestrian Detection

Segmentation / video segmentaion

DAVIS: Densely Annotated VIdeo Segmentation

SegTrack v2

Image Segmentation with A Bounding Box Prior dataset

Ground truth database of 50 images with: Data, Segmentation, Labelling - Lasso, Labelling - Rectangle

PASCAL VOC 2009 dataset

Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets

Motion Segmentation and OBJCUT data

Cows for object segmentation, Five video sequences for motion segmentation

Geometric Context Dataset

Geometric Context Dataset: pixel labels for seven geometric classes for 300 images

Crowd Segmentation Dataset

This dataset contains videos of crowds and other high density moving objects. The videos are collected mainly from the BBC Motion Gallery and Getty Images website. The videos are shared only for the research purposes. Please consult the terms and conditions of use of these videos from the respective websites.

CMU-Cornell iCoseg Dataset Contains hand-labelled pixel annotations for 38 groups of images, each group containing a common foreground. Approximately 17 images per group, 643 images total.

Segmentation evaluation database 200 gray level images along with ground truth segmentations

The Berkeley Segmentation Dataset and Benchmark

Image segmentation and boundary detection. Grayscale and color segmentations for 300 images, the images are divided into a training set of 200 images, and a test set of 100 images.

Weizmann horses

328 side-view color images of horses that were manually segmented. The images were randomly collected from the WWW.

Saliency-based video segmentation with sequentially updated priors

10 videos as inputs, and segmented image sequences as ground-truth

Daimler Urban Segmentation Dataset

The dataset consists of video sequences recorded in urban traffic. The dataset consists of 5000 rectified stereo image pairs. 500 frames come with pixel-level semantic class annotations into 5 classes: ground, building, vehicle, pedestrian, sky. Dense disparity maps are provided as a reference.

DAVIS: Densely Annotated VIdeo Segmentation 2016

A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation.

DAVIS: Densely Annotated VIdeo Segmentation 2017

The 2017 DAVIS Challenge on Video Object Segmentation.

CO-SKEL

Object Co-Skeletonization With Co-Segmentation.

CITY-OSM: Learning Aerial Image Segmentation From Online Maps

+20GB of aerial images obtained from Google Maps (including the groundtruth from OSM).

Mut1ny Face/Headsegmentation dataset

Head/face segmentation dataset contains over 16k labeled images.

The Unsupervised LLAMAS dataset

A lane marker detection and segmentation dataset of 100,000 images with 3d lines, pixel level dashed markers, and curves for individual lines.

Semantic labeling

Stanford background dataset

CamVid

Barcelona Dataset

SIFT Flow Dataset

References

https://pub.towardsai.net/50-object-detection-datasets-from-different-industry-domains-1a53342ae13d

Recognition / Facial recognition / Identification / Material Recognition / Scene Recognition / Fine-grained Visual Recognition

Face and Gesture Recognition Working Group FGnet

Face and Gesture Recognition Working Group FGnet

Feret

Face and Gesture Recognition Working Group FGnet

PUT face

9971 images of 100 people

Labeled Faces in the Wild

A database of face photographs designed for studying the problem of unconstrained face recognition

Urban scene recognition

Traffic Lights Recognition, Lara's public benchmarks.

PubFig: Public Figures Face Database

The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet. Unlike most other existing face datasets, these images are taken in completely uncontrolled situations with non-cooperative subjects.

YouTube Faces

The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is 48 frames, the longest clip is 6,070 frames, and the average length of a video clip is 181.3 frames.

MSRC-12: Kinect gesture data set

The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.

QMUL underGround Re-IDentification (GRID) Dataset

This dataset contains 250 pedestrian image pairs + 775 additional images captured in a busy underground station for the research on person re-identification.

Person identification in TV series

Face tracks, features and shot boundaries from our latest CVPR 2013 paper. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Big Bang Theory.

ChokePoint Dataset

ChokePoint is a video dataset designed for experiments in person identification/verification under real-world surveillance conditions. The dataset consists of 25 subjects (19 male and 6 female) in portal 1 and 29 subjects (23 male and 6 female) in portal 2.

Hieroglyph Dataset

Ancient Egyptian Hieroglyph Dataset.

Rijksmuseum Challenge Dataset: Visual Recognition for Art Dataset

Over 110,000 photographic reproductions of the artworks exhibited in the Rijksmuseum (Amsterdam, the Netherlands). Offers four automatic visual recognition challenges consisting of predicting the artist, type, material and creation year. Includes a set of baseline features, and offer a baseline based on state-of-the-art image features encoded with the Fisher vector.

The OU-ISIR Gait Database, Treadmill Dataset

Treadmill gait datasets composed of 34 subjects with 9 speed variations, 68 subjects with 68 subjects, and 185 subjects with various degrees of gait fluctuations.

The OU-ISIR Gait Database, Large Population Dataset

Large population gait datasets composed of 4,016 subjects.

Pedestrian Attribute Recognition At Far Distance

Large-scale PEdesTrian Attribute (PETA) dataset, covering more than 60 attributes (e.g. gender, age range, hair style, casual/formal) on 19000 images.

FaceScrub Face Dataset

The FaceScrub dataset is a real-world face dataset comprising 107,818 face images of 530 male and female celebrities detected in images retrieved from the Internet. The images are taken under real-world situations (uncontrolled conditions). Name and gender annotations of the faces are included.

Depth-Based Person Identification

Depth-Based Person Identification from Top View Dataset.

IMDb-Face: A large-scale noise-controlled face recognition dataset

IMDb-Face is a new large-scale noise-controlled dataset for face recognition research. The dataset contains about 1.7 million faces, 59k identities.

Open MIC dataset for Domain Adaptation and Few-shot Learning

ECCV 2018: Open Museum Identification Challenge dataset, photos of exhibits captured in 10 distinct exhibition spaces of several museums which showcase paintings, timepieces, sculptures, glassware, relics, science exhibits, natural history pieces, ceramics, pottery, tools and indigenous crafts.

Content-based image retrieval

CIFAR-10 2009

classes: 10
Training: 50,000
Test : 10,000
Image type: Object Category Images

NUS-WIDE 2009

classes: 21
Training : 97,214
Test: 65,075
Image type: Scene Images

MNIST 1998

classes : 10
Training: 60,000
Test : 10,000
Image type: Handwritten Digit Images

SVHN 2011

classes: 10
Training: 73,257
Test: 26,032
Image Type: House Number Images

SUN397 2010

classes: 397
Training: 100,754
Test: 8,000
Image type: Scene Images

UT-ZAP50K 2014

classes: 4
Training: 42,025
Test: 8,000
Image type: Shoes Images

Yahoo-1M 2015

classes: 116
Training: 1,011,723
Test: 112,363
Image type: Clothing Images

ILSVRC2012 2012

classes: 1,000
Training: ∼1.2 M
Test: 50,000
Image type: Object Category Images

MS COCO 2015

classes: 80
Training: 82,783
Test: 40,504
Image type: Common Object Images

MIRFlicker-1M 2010

Training : 1 M
Image type: Scene Images

Google Landmarks 2017

classes: 15 K
Training: ∼1 M
Image type: Landmark Images

Google Landmarks v2 2020

classes : 200 K
Training: 5 M
Image type: Landmark Images

Clickture 2013

Classes: 73.6 M
Training: 40 M
Image type: Search Log

Action(Action Classification in Video)/ Pose estimation / Human pose / Expression

Action (Action Classification in Video)

UCF Sports Action Dataset

This dataset consists of a set of actions collected from various sports which are typically featured on broadcast television channels such as the BBC and ESPN. The video sequences were obtained from a wide range of stock footage websites including BBC Motion gallery, and GettyImages.

UCF Aerial Action Dataset

This dataset features video sequences that were obtained using a R/C-controlled blimp equipped with an HD camera mounted on a gimbal.The collection represents a diverse pool of actions featured at different heights and aerial viewpoints. Multiple instances of each action were recorded at different flying altitudes which ranged from 400-450 feet and were performed by different actors.

UCF YouTube Action Dataset

It contains 11 action categories collected from YouTube.

Weizmann action recognition

Walk, Run, Jump, Gallop sideways, Bend, One-hand wave, Two-hands wave, Jump in place, Jumping Jack, Skip.

UCF50

UCF50 is an action recognition dataset with 50 action categories, consisting of realistic videos taken from YouTube.

ASLAN

The Action Similarity Labeling (ASLAN) Challenge.

MSR Action Recognition Datasets

The dataset was captured by a Kinect device. There are 12 dynamic American Sign Language (ASL) gestures, and 10 people. Each person performs each gesture 2-3 times.

KTH Recognition of human actions

Contains six types of human actions (walking, jogging, running, boxing, hand waving and hand clapping) performed several times by 25 subjects in four different scenarios: outdoors, outdoors with scale variation, outdoors with different clothes and indoors.

Hollywood-2 Human Actions and Scenes dataset

Hollywood-2 datset contains 12 classes of human actions and 10 classes of scenes distributed over 3669 video clips and approximately 20.1 hours of video in total.

Collective Activity Dataset

This dataset contains 5 different collective activities : crossing, walking, waiting, talking, and queueing and 44 short video sequences some of which were recorded by consumer hand-held digital camera with varying view point.

Olympic Sports Dataset

The Olympic Sports Dataset contains YouTube videos of athletes practicing different sports.

SDHA 2010

Surveillance-type videos

VIRAT Video Dataset

The dataset is designed to be realistic, natural and challenging for video surveillance domains in terms of its resolution, background clutter, diversity in scenes, and human activity/event categories than existing action recognition datasets.

HMDB: A Large Video Database for Human Motion Recognition

Collected from various sources, mostly from movies, and a small proportion from public databases, YouTube and Google videos. The dataset contains 6849 clips divided into 51 action categories, each containing a minimum of 101 clips.

Stanford 40 Actions Dataset

Dataset of 9,532 images of humans performing 40 different actions, annotated with bounding-boxes.

50Salads dataset

Fully annotated dataset of RGB-D video data and data from accelerometers attached to kitchen objects capturing 25 people preparing two mixed salads each (4.5h of annotated data). Annotated activities correspond to steps in the recipe and include phase (pre-/ core-/ post) and the ingredient acted upon.

Penn Sports Action The dataset contains 2326 video sequences of 15 different sport actions and human body joint annotations for all sequences.

CVRR-HANDS 3D

A Kinect dataset for hand detection in naturalistic driving settings as well as a challenging 19 dynamic hand gesture recognition dataset for human machine interfaces.

TUM Kitchen Data Set

Observations of several subjects setting a table in different ways. Contains videos, motion capture data, RFID tag readings,...

TUM Breakfast Actions Dataset

This dataset comprises of 10 actions related to breakfast preparation, performed by 52 different individuals in 18 different kitchens.

MPII Cooking Activities Dataset

Cooking Activities dataset.

GTEA Gaze+ Dataset

This dataset consists of seven meal-preparation activities, each performed by 10 subjects. Subjects perform the activities based on the given cooking recipes.

UTD-MHAD: multimodal human action recogniton dataset

The dataset consists of four temporally synchronized data modalities. These modalities include RGB videos, depth videos, skeleton positions, and inertial signals (3-axis acceleration and 3-axis angular velocity) from a Kinect RGB-D camera and a wearable inertial sensor for a comprehensive set of 27 human actions.

Action Recognition Datasets: "NTU RGB+D" Dataset and "NTU RGB+D 120" Dataset

"NTU RGB+D" contains 60 action classes and 56,880 video samples. "NTU RGB+D 120" extends "NTU RGB+D" by adding another 60 classes and another 57,600 video samples, i.e., "NTU RGB+D 120" has 120 classes and 114,480 samples in total. These two datasets both contain RGB videos, depth map sequences, 3D skeletal data, and infrared (IR) videos for each sample. Each dataset is captured by three Kinect V2 cameras concurrently.

Pose estimation / Human pose / Expression

Pose Estimation is a computer vision technique to predict and track the location of a person or object.
This is typically done by identifying, locating, and tracking a number of keypoints on a given object or person.

Leeds Sport Poses

The LSP dataset contains 10,000 images gathered from Flickr searches for the tags "parkour", "gymnastics", and "athletics".

AFEW (Acted Facial Expressions In The Wild)/SFEW (Static Facial Expressions In The Wild)

Dynamic temporal facial expressions data corpus consisting of close to real world environment extracted from movies.

Expression in-the-Wild (ExpW) Dataset

Contains 91,793 faces manually labeled with expressions. Each of the face images was manually annotated as one of the seven basic expression categories: â€œangryâ€�, â€œdisgustâ€�, â€œfearâ€�, â€œhappyâ€�, â€œsadâ€�, â€œsurpriseâ€�, or â€œneutralâ€�.

ETHZ CALVIN Dataset

CALVIN research group datasets

HandNet (annotated depth images of articulating hands)

This dataset includes 214971 annotated depth images of hands captured by a RealSense RGBD sensor of hand poses. Annotations: per pixel classes, 6D fingertip pose, heatmap. Images -> Train: 202198, Test: 10000, Validation: 2773. Recorded at GIP Lab, Technion.

3D Human Pose Estimation

Depth videos + ground truth human poses from 2 viewpoints to improve 3D human pose estimation.

Dynamic Faust

More than 40.000 scans of people very accurately registered. Scans contain texture so synthetic videos/images are easy to generate. See also Dyna: A Model of Dynamic Human Shape in Motio.

BUFF dataset

About 10.000 scans of people in clothing and the estimated body shape of people underneath. Scans contain texture so synthetic videos/images are easy to generate.

TNT 15 dataset

Several sequences of video synchronised by 10 Inertial Sensors (IMU) worn at the extremities.

Extended Chictopia dataset

Chictopia dataset with additional processed annotations (face) and SMPL body model fits to the images. The copyright of the images belongs to the original authors of Chictopia.

Professional Portrait Dataset

Over 320,000 highly-rated portrait images crawled from 500px website.

Optical character recognition (OCR)

Text OCR

28,134 natural images from TextVQA
903,069 annotated scene-text words
32 words per image on average

NIST Database

The US National Institute of Science publishes handwriting from 3600 writers, including more than 800,000 character images.

FUNSD

Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.

ICDAR 2003

The ICDAR2003 dataset is a dataset for scene text recognition. It contains 507 natural scene images (including 258 training images and 249 test images) in total.

ST-VQA

ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.

Devangri Characters

A dataset of handwritten Devangari characters, composed of 1800 samples from 36 character classes obtained by 25 native writers.

Mathematics Expressions

More than 10,000 expressions, including more than 101 mathematical symbols.

Chinese Characters

A dataset of handwritten Chinese characters containing 909,818 images that corresponds to about 10 news articles.

Arabic Printed Text

Contains a lexicon of 113,284 words, and uses 10 Arabic fonts.

Document database

Contains 941 online handwritten documents by 189 writers, and covers lists, tables, formulas, diagrams and drawings.

Iam On-line Handwriting

Contains forms of handwritten English text acquired on a whiteboard, and includes more than 1700 entries.

Street View Text

The Street View Text dataset was harvested from Google Street View, and mostly deals with outdoor street level signs and boards.

Street View House Numbers

Contains 73257 digits of house street numbers, taken from Google Street View.

Natural Environment OCR

A dataset that contains 659 real world images with 5238 annotations of text.

Scene Text

Contains 3000 images captured in different environments, including outdoors and indoors scenes under different lighting conditions (clear day, night, strong artificial lights, etc).

Text Detection

Contains 500 natural images, which are taken using a pocket camera. The indoor images are mainly signs, doorplates and caution plates while the outdoor images are mostly guide boards and billboards.

Stanford OCR

Contains handwritten words dataset collected by MIT Spoken Language Systems Group, published by Stanford.

Chars74K Data

This has 74K images of both English and Kannada digits.

2D code reading

2D code reading – reading of 2D codes such as data matrix and QR codes.

QR-DN1.0 dataset

The QR-DN1.0 dataset includes 5 categories of QR codes that will cover low to high density levels. Each group has 15 QR codes: 5 images for testing and 10 images for training.

Dynamosoft

The data set has 57 valid images with 112 Data Matrix codes. The images are put into three categories: synthetic, Internet and industrial.

Shape Recognition Technology (SRT)

SHAPES

There are 244 questions and 15,616 images in total, with all questions having a yes and no answer (and corresponding supporting image). Each image is a 30×30 RGB image depicting a 3×3 grid of objects. Each object is characterized by shape (circle, square, triangle), colour (red, green, blue) and size (small, big).

Egomotion

Egomotion refers to estimating a camera's motion relative to a rigid scene.

Stereo Ego-Motion Dataset

This dataset contains 494 full-HD videos across 4 categories - car, cat, chair, dog.

Paired Egocentric Video (PEV) Dataset (CVPR’16)

video tracking/ object tracking

BIWI Walking Pedestrians dataset

Walking pedestrians in busy scenarios from a bird eye view

"Central" Pedestrian Crossing Sequences

Three pedestrian crossing sequences

Pedestrian Mobile Scene Analysis

The set was recorded in Zurich, using a pair of cameras mounted on a mobile platform. It contains 12'298 annotated pedestrians in roughly 2'000 frames.

Head tracking

BMP image sequences.

KIT AIS Dataset

Data sets for tracking vehicles and people in aerial image sequences.

MIT Traffic Data Set

MIT traffic data set is for research on activity analysis and crowded scenes. It includes a traffic video sequence of 90 minutes long. It is recorded by a stationary camera.

Shinpuhkan 2014 dataset: Multi-Camera Pedestrian Dataset for Tracking People across Multiple Cameras

This dataset consists of more than 22,000 images of 24 people which are captured by 16 cameras installed in a shopping mall "Shinpuh-kan". All images are manually cropped and resized to 48x128 pixels, grouped into tracklets and added annotation.

ATC shopping center dataaset

The tracking environment consists of multiple 3D range sensors, covering an area of about 900 m2, in the "ATC" shopping center in Osaka, Japan.

Human detection and tracking using RGB-D camera

Collected in a clothing store. Captured with Kinect (640*480, about 30fps)

Multiple Camera Tracking

Hallway Corridor - Multiple Camera Tracking: An indoor camera network dataset with 6 cameras (contains ground plane homography).

Multiple Object Tracking Benchmark

A centralized benchmark for multi-object tracking.

Stanford Drone Dataset

The dataset consists of eight unique scenes in crowded spaces such as a university campus or the sidewalks of a busy street.

Moving cameras monitoring the same scenes

Low-resolution RGB videos + ground truth trajectories from multiple fixed and moving cameras monitoring the same scenes (indoor and outdoor) to improve object tracking and matching.

Visual Tracking

Visual Tracker Benchmark

Visual Tracker Benchmark v1.1

VOT Challenge

Princeton Tracking Benchmark

Tracking Manipulation Tasks (TMT)

Optical flow

To determine, for each point in the image, how that point is moving relative to the image plane, i.e., its apparent motion. This motion is a result both of how the corresponding 3D point is moving in the scene and how the camera is moving relative to the scene.

Middlebury Optical Flow Evaluation

MPI-Sintel Optical Flow Dataset and Evaluation

The KITTI Vision Benchmark Suite

HCI Challenge

Scene reconstruction / multiview(multiview reconstruction) / 3D Face Animation

3D scene data

Stanford 3D Scene Data

Multi-View Stereo Reconstruction

RetrievalFuse

3D Photography Dataset

Multiview stereo data sets: a set of images

Multi-view Visual Geometry group's data set

Dinosaur, Model House, Corridor, Aerial views, Valbonne Church, Raglan Castle, Kapel sequence

Oxford reconstruction data set (building reconstruction)

Oxford colleges

Multi-View Stereo dataset (Vision Middlebury)

Temple, Dino

Multi-View Stereo for Community Photo Collections

Venus de Milo, Duomo in Pisa, Notre Dame de Paris

IS-3D Data

Dataset provided by Center for Machine Perception

CVLab dataset

CVLab dense multi-view stereo image database

3D Objects on Turntable

Objects viewed from 144 calibrated viewpoints under 3 different lighting conditions

Object Recognition in Probabilistic 3D Scenes

Images from 19 sites collected from a helicopter flying around Providence, RI. USA. The imagery contains approximately a full circle around each site.

Multiple cameras fall dataset

24 scenarios recorded with 8 IP video cameras. The first 22 first scenarios contain a fall and confounding events, the last 2 ones contain only confounding events.

CMP Extreme View Dataset

15 wide baseline stereo image pairs with large viewpoint change, provided ground truth homographies.

KTH Multiview Football Dataset II

This dataset consists of 8000+ images of professional footballers during a match of the Allsvenskan league. It consists of two parts: one with ground truth pose in 2D and one with ground truth pose in both 2D and 3D.

Disney Research light field datasets

This dataset includes: camera calibration information, raw input images we have captured, radially undistorted, rectified, and cropped images, depth maps resulting from our reconstruction and propagation algorithm, depth maps computed at each available view by the reconstruction algorithm without the propagation applied.

CMU Panoptic Studio Dataset

Multiple people social interaction dataset captured by 500+ synchronized video cameras, with 3D full body skeletons and calibration data.

4D Light Field Dataset

24 synthetic scenes. Available data per scene: 9x9 input images (512x512x3) , ground truth (disparity and depth), camera parameters, disparity ranges, evaluation masks.

3D Face Animation

Image Generation/Image restoration / De-noising / Super-Resolution

Super-Resolution

Single-Image Super-Resolution: A Benchmark

Image Deblurring

Sun dataset

Levin dataset

Stereo Vision

Is the process of extracting 3D information from digital images taken by two cameras displaced horizontally from one another to obtain two different views of the same scene.

Middlebury Stereo Vision

The KITTI Vision Benchmark Suite

LIBELAS: Library for Efficient Large-scale Stereo Matching

Ground Truth Stixel Dataset

Intrinsic Images

Ground-truth dataset and baseline evaluations for intrinsic image algorithms

Intrinsic Images in the Wild

Intrinsic Image Evaluation on Synthetic Complex Scenes

Visual Surveillance

VIRAT

CAM2

CAVIAR

For the CAVIAR project a number of video clips were recorded acting out the different scenarios of interest. These include people walking alone, meeting with others, window shopping, entering and exitting shops, fighting and passing out and last, but not least, leaving a package in a public place.

ViSOR

ViSOR contains a large set of multimedia data and the corresponding annotations.

CUHK Crowd Dataset

474 video clips from 215 crowded scenes, with ground truth on group detection and video classes.?

TImes Square Intersection (TISI) Dataset

A busy outdoor dataset for research on visual surveillance.

Educational Resource Centre (ERCe) Dataset

An indoor dataset collected from a university campus for physical event understanding of long video streams.

PIROPO Database: People in Indoor ROoms with Perspective and Omnidirectional cameras

Multiple sequences recorded in two different indoor rooms, using both omnidirectional and perspective cameras, containing people in a variety of situations (people walking, standing, and sitting). Both annotated and non-annotated sequences are provided, where ground truth is point-based. In total, more than 100,000 annotated frames are available.

UNICITY: A depth maps database for people detection in security airlocks

UNICITY consists of 58k images collected from 65 recorded sequences with one or two people performing different behaviors including attacks and trickeries. It also provides full annotation of people such as the location of head and shoulders.

Image Captioning

Flickr 8K

Flickr 30K

Microsoft COCO

Foreground/Background

Wallflower Dataset

For evaluating background modelling algorithms

Foreground/Background Microsoft Cambridge Dataset

Foreground/Background segmentation and Stereo dataset from Microsoft Cambridge

Stuttgart Artificial Background Subtraction Dataset

The SABS (Stuttgart Artificial Background Subtraction) dataset is an artificial dataset for pixel-wise evaluation of background models.

Image Alpha Matting Dataset

Image Alpha Matting Dataset.

LASIESTA: Labeled and Annotated Sequences for Integral Evaluation of SegmenTation Algorithms

LASIESTA is composed by many real indoor and outdoor sequences organized in diferent categories, each of one covering a specific challenge in moving object detection strategies.

Image stitching

IPM Vision Group Image Stitching datasets

Images and parameters for registeration

Medical

VIP Laparoscopic / Endoscopic Dataset

Collection of endoscopic and laparoscopic (mono/stereo) videos and images

Mouse Embryo Tracking Database

DB Contains 100 examples with the uncompressed frames, up to the 10th frame after the appearance of the 8th cell; a text file with the trajectories of all the cells, from appearance to division; a movie file showing the trajectories of the cells.

FIRE Fundus Image Registration Dataset

134 retinal image pairs and ground truth for registration.

deepchand41 / motion-planning Goto Github PK

motion-planning's Introduction

computer vision dataset

Contributing and Collaborating

Object recognition (also called object classification)

Image Classification

Medical Images

Agriculture and Scene

Video Classification

Automobile and ADAS Related Datasets

References

Detection(object detection,Edge detection)/ Multi-view Object Detection / Segmentation(Image Segmentation) / Saliency(salient) Detection / Semantic labeling

Detection

Multi-view Object Detection

Saliency Detection

Pedestrian Detection

Segmentation / video segmentaion

Semantic labeling

References

Recognition / Facial recognition / Identification / Material Recognition / Scene Recognition / Fine-grained Visual Recognition

Material Recognition

Scene Recognition

Fine-grained Visual Recognition

Content-based image retrieval

Action(Action Classification in Video)/ Pose estimation / Human pose / Expression

Action (Action Classification in Video)

Pose estimation / Human pose / Expression

Optical character recognition (OCR)

2D code reading

Shape Recognition Technology (SRT)

Egomotion

video tracking/ object tracking

Optical flow

Scene reconstruction / multiview(multiview reconstruction) / 3D Face Animation

3D Face Animation

Image Generation/Image restoration / De-noising / Super-Resolution

Super-Resolution

Image Deblurring

Stereo Vision

Intrinsic Images

Visual Surveillance

Image Captioning

Foreground/Background

Image stitching

Medical

References

Recommend Projects

Recommend Topics

Recommend Org

Jobs