Search references for DATASET SHIFT. Phrases containing DATASET SHIFT
See searches and references containing DATASET SHIFT!DATASET SHIFT
Change in data distribution between training and deployment in machine learning
Dataset shift is a phenomenon in machine learning and statistics in which the joint distribution of input variables and target labels is different in
Dataset_shift
These datasets are used in machine learning (ML) research and have been cited in peer-reviewed academic journals. Datasets are an integral part of the
List of datasets for machine-learning research
List_of_datasets_for_machine-learning_research
Series of language models developed by Google AI
reason not all selected tokens are masked is to avoid the dataset shift problem. The dataset shift problem arises when the distribution of inputs seen during
BERT_(language_model)
Database of handwritten digits
original datasets. The creators felt that since NIST's training dataset was taken from American Census Bureau employees, while the testing dataset was taken
MNIST_database
Change in wavelength of light
"The 2dF galaxy redshift survey: Power-spectrum analysis of the final dataset and cosmological implications". Monthly Notices of the Royal Astronomical
Redshift
This is a list of datasets for machine learning research. It is part of the list of datasets for machine-learning research. These datasets consist primarily
List of datasets in computer vision and image processing
List_of_datasets_in_computer_vision_and_image_processing
Tasks in machine learning
ISBN 978-3-642-35289-8. "Machine learning - Is there a rule-of-thumb for how to divide a dataset into training and validation sets?". Stack Overflow. Retrieved 2021-08-12
Training, validation, and test data sets
Training,_validation,_and_test_data_sets
Artificial intelligence field of study
Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift". NeurIPS. arXiv:1906.02530. Bogdoll, Daniel; Breitenstein, Jasmin;
AI_safety
Image dataset
The CIFAR-10 dataset (Canadian Institute For Advanced Research) is a collection of images that are commonly used to train machine learning and computer
CIFAR-10
Type of large language model
learning architecture called the transformer. They are pre-trained on large datasets of unlabeled content, and able to generate novel content. OpenAI was the
Generative pre-trained transformer
Generative_pre-trained_transformer
Type of machine learning model
question. Some datasets are adversarial, focusing on problems that confound LLMs. One example is the TruthfulQA dataset, a question answering dataset consisting
Large_language_model
Method in machine learning
dataset. The original dataset is whatever information is given. The bootstrap dataset is made by randomly picking objects from the original dataset.
Bootstrap_aggregating
Mathematical technique
Mean shift is a non-parametric feature-space mathematical analysis technique for locating the maxima of a density function, a so-called mode-seeking algorithm
Mean_shift
2018 text-generating language model
labeled data. This reliance on supervised learning limited their use of datasets that were not well-annotated, in addition to making it prohibitively expensive
GPT-1
Vector quantization algorithm minimizing the sum of squared deviations
adopted in early machine learning and data analysis tasks involving large datasets. Despite its widespread use, limitations such as sensitivity to initial
K-means_clustering
Class of nonparametric methods
Covariate shift and local learning by distribution matching. In J. Quinonero-Candela, M. Sugiyama, A. Schwaighofer, N. Lawrence (eds.). Dataset shift in machine
Kernel embedding of distributions
Kernel_embedding_of_distributions
Landsat-based dataset of global tree cover change
2000–2012. This shift from a landmark study to a regularly updated data product made Global Forest Change a long-lived reference dataset for later forest-monitoring
Global_Forest_Change_dataset
Field associated with machine learning and transfer learning
labeled in another. Prior Shift (Label Shift) occurs when the label distribution differs between the source and target datasets, while the conditional distribution
Domain_adaptation
Restrictions on using GPS in China
conversion methods both ways largely renders obsolete datasets for deviations mentioned below. The China GPS shift (or offset) problem is a class of issues stemming
Restrictions on geographic data in China
Restrictions_on_geographic_data_in_China
Process supporting machine learning
Data annotation is the process within a dataset of adding relevant metadata labels or tags to enable machines to interpret the data in line with its intended
Data_annotation
Data visualization
box-and-whisker diagram. Outliers that differ significantly from the rest of the dataset may be plotted as individual points beyond the whiskers on the box plot
Box_plot
Type of artificial intelligence system
Training used three datasets: The LTIP dataset mentioned above, a curated dataset of video-text pairs (called VTP), and a massive dataset of interleaved text-image
Vision-language_model
Evolution of a word's meaning
Google Books n-gram dataset and the Corpus of Historical American English (COHA). Database of Semantic Shifts in languages of the world (DatSemShift 3.0)
Semantic_change
Type of feedforward neural network
etc.) Robust datasets also increase the probability that CNNs will learn the generalized principles that characterize a given dataset rather than the
Convolutional_neural_network
the Global Study on Homicide are based on the UNODC Homicide Statistics dataset, which is derived from the criminal justice or public health systems of
List of countries by intentional homicide rate
List_of_countries_by_intentional_homicide_rate
Measure of similarity between two data clusterings
Example of two similar clusterings of the same dataset, produced by k-means and mean shift. The adjusted Rand index between them is A R I ≈ 0.94 {\displaystyle
Rand_index
Concept in machine learning
entries across dataset splits is also important. For language models, the Min-K% method can detect the presence of data in a pretraining dataset. It presents
Leakage_(machine_learning)
Neural network that learns efficient data encoding in an unsupervised manner
the reference distribution is just the empirical distribution given by a dataset { x 1 , . . . , x N } ⊂ X {\displaystyle \{x_{1},...,x_{N}\}\subset {\mathcal
Autoencoder
Loss-of-control incident at OpenAI
models, datasets, and demonstration applications, processing user-uploaded content such as model weights and datasets. Some supported dataset formats
OpenAI–HuggingFace_incident
2023 text-generating language model
code models. Granite models are trained on datasets curated from Internet, academic publishings, code datasets, legal and finance documents. A foundation
IBM_Granite
Approach in generative models
of a target dataset and generates a similar but larger dataset. EBMs detect the latent variables of a dataset and generate new datasets with a similar
Energy-based_model
Technique in neural networks for learning joint representations of text and images
To train a pair of CLIP models, one would start by preparing a large dataset of image-caption pairs. During training, the models are presented with
Contrastive Language–Image Pre-training
Contrastive_Language–Image_Pre-training
Artificial intelligence research collective
to GPT-3. On December 31, 2020, EleutherAI released The Pile, a curated dataset of diverse text for training large language models. While the paper referenced
EleutherAI
Flaw in mathematical modelling
fitted relationship will appear to perform less well on a new dataset than on the dataset used for fitting (a phenomenon sometimes known as shrinkage)
Overfitting
error and possibly confirmation bias, although modern reanalysis of the dataset suggests that Eddington's analysis was accurate. The measurement was repeated
Tests_of_general_relativity
Artificial intelligence model paradigm
sound, etc.), is a machine learning or deep learning model trained on vast datasets so that it can be applied across a wide range of use cases. Generative
Foundation_model
Technique used to record and analyze human movement
by advances in computer vision and the availability of several training datasets. These systems can estimate full-body pose in two or three dimensions,
Markerless_motion_capture
American multinational technology conglomerate
2005. In 2021, it rebranded as Meta Platforms, Inc. to reflect a strategic shift toward developing the metaverse—an interconnected digital ecosystem spanning
Meta_Platforms
Machine learning technique
collection models, where the model is learning by interacting with a static dataset and updating its policy in batches, as well as online data collection models
Reinforcement learning from human feedback
Reinforcement_learning_from_human_feedback
Algorithm for supervised learning of binary classifiers
is proved by Rosenblatt et al. Perceptron convergence theorem—Given a dataset D {\textstyle D} , such that max ( x , y ) ∈ D ‖ x ‖ 2 = R {\textstyle
Perceptron
Machine learning methods using multiple input modalities
and image tokens. The compound model is then fine-tuned on an image-text dataset. This basic construction can be applied with more sophistication to improve
Multimodal_learning
Iterative method for finding maximum likelihood estimates in statistical models
algorithm fitting a two component Gaussian mixture model to the Old Faithful dataset. The algorithm steps through from a random initialization to convergence
Expectation–maximization algorithm
Expectation–maximization_algorithm
Deep learning architecture
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Mamba (deep learning architecture)
Mamba_(deep_learning_architecture)
Tree-based ensemble machine learning methods
\mathbf {x} } , designed with randomness Θ j {\displaystyle \Theta _{j}} and dataset D n {\displaystyle {\mathcal {D}}_{n}} , and N n ( x , Θ j ) = ∑ i = 1
Random_forest
Statistical hypothesis test
a t-test because the latter converges to the former as the size of the dataset increases. The term "t-statistic" is abbreviated from "hypothesis test
Student's_t-test
Density-based data clustering algorithm
DBSCAN can find non-linearly separable clusters. This dataset cannot be adequately clustered with k-means or Gaussian Mixture EM clustering.
DBSCAN
Computational model used in machine learning
accelerated by the use of graphics processing units (GPUs), and large datasets. Simplified example of training a neural network in object detection: The
Neural network (machine learning)
Neural_network_(machine_learning)
Conversational software
responses word by word based on user input, and are usually trained on a large dataset of natural-language phrases. They sometimes provide plausible-sounding
Chatbot
Problem setup in machine learning
03228. Yin, Wenpeng (2019). "Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach" (PDF). EMNLP. arXiv:1909.00161. Levy
Zero-shot_learning
Paradigm in machine learning that uses no classification labels
data, training, algorithm, and downstream applications. Typically, the dataset is harvested cheaply "in the wild", such as massive text corpus obtained
Unsupervised_learning
Machine learning framework
known dataset serves as the initial training data for the discriminator. Training involves presenting it with samples from the training dataset until
Generative adversarial network
Generative_adversarial_network
Algorithm for modelling sequential data
adopted for training large language models (LLMs) on large (language) datasets. Modern transformer designs are commonly grouped into encoder-only, decoder-only
Transformer_(deep_learning)
Method of measuring prediction error
samples and OOB sets are created. The OOB sets can be aggregated into one dataset, but each sample is only considered out-of-bag for the trees that do not
Out-of-bag_error
Ensemble learning method
BrownBoost, can learn from noisy datasets and can specifically learn the underlying classifier of the Long–Servedio dataset. Random forest Alternating decision
Boosting_(machine_learning)
Neural network technology
CNN architecture for handwritten digit recognition, trained on the MNIST dataset, and was used in ATM. (Olshausen & Field, 1996) discovered that simple
Convolutional_layer
Human-caused changes to climate on Earth
Archived from the original on 13 January 2026. Click on "Datasets". Natural driver dataset is downloadable by clicking on "gmst_changes_model_and_obs
Climate_change
Modem for computers released by AT&T in 1962
followed the introduction of the 110 baud Bell 101 dataset in 1958. The Bell 103 modem used audio frequency-shift keying to encode data. Different pairs of audio
Bell_103
Statistical model of language
as of 2026, are predominantly based on transformers trained on larger datasets (frequently using texts scraped from the public internet). They have superseded
Language_model
Effect of variables' uncertainties on the uncertainty of a function based on them
sampling techniques from the Monte Carlo method family. For very large datasets or complex functions, the calculation of the error propagation may be very
Propagation_of_uncertainty
Number of distinct species in a biological community
number of different species that are represented in a given community (a dataset). The effective number of species refers to the number of equally abundant
Species_diversity
Deep learning generative model to encode data representation
Thus, the encoder maps each point (such as an image) from a large complex dataset into a distribution within the latent space, rather than to a single point
Variational_autoencoder
2019 text-generating language model
in their foundational series of GPT models. GPT-2 was pre-trained on a dataset of 8 million web pages. It was partially released in February 2019, followed
GPT-2
Country in West Asia
Bank Open Data". World Bank Open Data. Retrieved 10 March 2025. "Iran Datasets". IMF. Retrieved 10 March 2025. Wehrey, Frederic; Green, Jerrold D.; Nichiporuk
Iran
3D reconstruction technique
require a specialized camera or software. Any camera is able to generate datasets, provided the settings and capture method meet the requirements for SfM
Neural_radiance_field
Set of learning techniques in machine learning
unlabeled input data by analyzing the relationship between points in the dataset. Examples include dictionary learning, independent component analysis,
Representation_learning
Influential 2012 deep convolutional neural network
performing over 2,200 forward passes per second under ideal conditions. The dataset images were stored in JPEG format. They took up 27GB of disk. The neural
AlexNet
Important algorithms in numerical statistics
print(ab.get_sample_variance()) This algorithm allows one to divide a dataset into several pieces, run them in parallel, and then merge the results together
Algorithms for calculating variance
Algorithms_for_calculating_variance
Class of algorithms for pattern analysis
clusters, rankings, principal components, correlations, classifications) in datasets. For many algorithms that solve these tasks, the data in raw representation
Kernel_method
Machine learning calibration technique
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Platt_scaling
Machine learning technique
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Mixture_of_experts
Mathematical technique in spectroscopic analysis
adapted accordingly. Suppose the original dataset D contains the n spectra in rows. The signals of the original dataset are generally preprocessed. The original
Two-dimensional correlation analysis
Two-dimensional_correlation_analysis
Machine learning-powered structure design
Barret Zoph and Quoc Viet Le applied NAS with RL targeting the CIFAR-10 dataset and achieved a network architecture that rivals the best manually-designed
Neural_architecture_search
Government digital portal in Saudi Arabia
Artificial Intelligence Authority (SDAIA). It provides public access to datasets released by various government entities in standardized, reusable formats
Open.data.gov.sa
Extension of cubic spline interpolation
f_{x}} from those. The two give equivalent results.) At the edges of the dataset, when one is missing some of the surrounding points, the missing points
Bicubic_interpolation
Integrated circuit technology
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Neuromorphic_computing
Tornado research experiment
longetivity using in-situ probes that gathered thermodynamic and kinematic datasets at close proximity to tornadoes, and mobile mesonet stations that were
TWISTEX
Checking software against expectations
needed. Test development: test procedures, test scenarios, test cases, test datasets, test scripts to use in testing software. Test execution: testers execute
Software_testing
Type of feedforward neural network
receives the initial data signal (such as pixels of an image or features of a dataset) and passes it to the first hidden layer. It does not perform any computation
Multilayer_perceptron
Insincere flattery, once meant a false accuser
'informer'), and Italian. In modern English, the meaning of the word has shifted to mean flattery. The origin of the Ancient Greek word συκοφάντης (sykophántēs)
Sycophancy
Technique for the generative modeling of a continuous probability distribution
process for a given dataset, such that the process can generate new elements that are distributed similarly as the original dataset. A diffusion model
Diffusion_model
Subset of artificial intelligence
partition a dataset into a specified number of clusters, k, each represented by the centroid of its points. This process condenses extensive datasets into a
Machine_learning
Similarity measure for number sequences
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Cosine_similarity
OpenAI text-to-video models (2024–2026)
Brooks stated that the model learned how to create 3D graphics from its dataset alone, while fellow Sora researcher Bill Peebles said that the model automatically
Sora_(text-to-video_model)
1986 nuclear accident in the Soviet Union
ISBN 978-5-93728-006-0. OCLC 539577824. EBSCOhost XISI321014-H (from dataset "ISIS Current Bibliography of History of Science"). See alternatively 2005
Chernobyl_disaster
Software user interface
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Human-in-the-loop
Technique in machine learning
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Curriculum_learning
2023 text-generating language model
given large datasets of text taken from the internet and trained to predict the next token (roughly corresponding to a word) in those datasets. Second, human
GPT-4
Machine learning technique
In 1981, a report considered the application of transfer learning to a dataset of images representing letters of computer terminals, experimentally demonstrating
Transfer_learning
Extracting features from raw data for machine learning
engineering has been clustering of feature-objects or sample-objects in a dataset. Especially, feature engineering based on matrix decomposition has been
Feature_engineering
Measures of observational error
it tends to be greatly affected by the particular class prevalence in a dataset and the classifier's biases. Furthermore, it is also called top-1 accuracy
Accuracy_and_precision
Act of violence committed in support of environmental causes
until the mid-2000s when it increasingly shifted to civil disobedience and mass protests. Large-N datasets have found very few instances of fatalities
Eco-terrorism
Method used to normalize the range of independent variables
Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift". arXiv:1502.03167 [cs.LG]. Juszczak, P.; D. M. J. Tax; R. P. W. Dui (2002)
Feature_scaling
Process of analyzing large data sets
mining is that data analysis is used to test models and hypotheses on the dataset, e.g., analyzing the effectiveness of a marketing campaign, regardless
Data_mining
Family of large language models by Alibaba
the Apache 2.0 License, although only the weights were released, not the dataset or training method. QwQ has a 32K token context length and performs better
Qwen
Reverse-engineering neural networks
articles Glossary of artificial intelligence List of datasets for machine-learning research List of datasets in computer vision and image processing Outline
Mechanistic_interpretability
Machine learning technique
Normalization techniques are often theoretically justified as reducing covariance shift, smoothing optimization landscapes, and increasing regularization, though
Normalization (machine learning)
Normalization_(machine_learning)
Interatomic potentials constructed by machine learning programs
structures of a single material. A major shift occurred with the creation of large, chemically diverse datasets enabling models that generalize across many
Machine-learned interatomic potential
Machine-learned_interatomic_potential
Shadow library search engine
WorldCat, and Google Books are listed as metadata-only sources. Some of these datasets are already publicly accessible, while others are scraped or otherwise
Anna's_Archive
Statistics and machine learning technique
the output of each individual classifier or regressor for the entire dataset can be viewed as a point in a multi-dimensional space. Additionally, the
Ensemble_learning
Method of improving artificial neural network
normalization works so well. It was initially thought to tackle internal covariate shift, a problem where parameter initialization and changes in the distribution
Batch_normalization
travel, tourism, insurance
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
DATASET SHIFT
travel, tourism, insurance