HOW CAN MACHINES LEARN? –DATASET
Source: https://pub.towardsai.net/best–datasets-for-machine-learning-data-science-computer–
vision–nlp–ai-c9541058cf4f
Computer Vision Datasets
xView:xView is one of the most
massive publicly available datasets of
overhead imagery. It contains images
from complex scenes around the
world, annotated using bounding
boxes.
ImageNet: The largest image
dataset for computer vision. It
provides an accessible image
database that is organized
hierarchically, according to WordNet.
Kinetics-700:A large-scale dataset
of video URLs from Youtube.
Including human-centered actions. It
contains over 700,000 videos.
Google’s Open Images:A vast
dataset from Google AI containing
over 10 million images.
Sentiment Analysis Datasets
Lexicoder Sentiment Dictionary:This dataset is
specific for sentiment analysis. The dataset
contains over 3000 negative words and over
2000 positive sentiment words.
IMDB reviews: An interesting dataset with
over 50,000 movie reviews from Kaggle.
Stanford Sentiment Treebank: Standard
sentiment dataset with sentiment annotations.
Twitter US Airline Sentiment:Twitter data on
US airlines from February 2015, classified as
positive, negative, and neutral tweets
Self-driving (Autonomous Driving) Datasets
Waymo Open Dataset:This is a fantastic
dataset resource from the folks at Waymo.
Includes a vast dataset of autonomous driving,
enough to train deep nets from zero.
Berkeley DeepDrive BDD100k:One of the
largest datasets for self-driving cars, containing
over 2000 hours of driving experiences across
New York and California.
Bosch Small Traffic Light Dataset: Dataset for
small traffic lights for deep learning.
Clinical Datasets
MaskedFace-Net:MaskedFace-Net is a real
dataset containing human faces with correct
and incorrectly worn masks. It contains over
137k images which are based on the Flick–
Faces–HQ dataset [21]. For more details about
the dataset and its uses, please visit
the documentation on Github.
COVID-19 Dataset: The Allen Institute of AI
research has released a vast research dataset
of over 45,000 scholarly articles about COVID–
19.
MIMIC-III: Openly available dataset developed
by the MIT Lab for Computational Physiology,
comprising de-identified health data
associated with ~40,000 critical care patients.
It includes demographics, vital signs,
laboratory tests, medications, and more.