TANTI TECHNOLOGIES
Welcome to My Technology Blog Hi, I’m Sandeep Kumar, Founder of Tanti Technologies and a Data Scientist & Technology Leader with 10+ years of experience in Data, Cloud, AI, DevOps, AIX, and Linux. Through this blog, I share practical insights, technical guides, real-world solutions, and industry knowledge covering Data Engineering, Cloud, DevOps, AI/ML, GenAI, and Agentic AI. My goal is to simplify complex technologies and help professionals learn, build, automate, and innovate.
Tanti Technology
- sandeep tanti
- Bangalore, karnataka, India
- Sandeep Kumar | Founder, Tanti Technologies At Tanti Technologies, I share practical knowledge, hands-on tutorials, and real-world solutions across Python, SQL, Docker, Kubernetes, CI/CD, AWS, Azure, Machine Learning, Generative AI, and Agentic AI. My mission is to simplify complex technologies and provide actionable insights that help developers, engineers, and businesses learn faster, build better, automate smarter, and innovate with confidence. š Learn • Build • Automate • Innovate Follow Tanti Technologies for practical technology insights and next-generation AI solutions.
Wednesday, 30 September 2026
Monday, 28 September 2026
5. Dataset and subsets:
Training Set:
·
Training set is used for learning
the parameters of the model.
· Training set has an optimistic bias.
· The training set is usually
the biggest one of three sets and used to build the model.
Test Set:
·
Test set is used to get a final, unbiased estimate
of how well the learning
method works.
·
We expect that this estimate to be worse than training/validation set.
· Test set is not biased
and this distinguishes test set from training
set.
· Similar to Training set, Test set is a finite
sample and bound to have variance due to sample size, but the test
set doesn’t have an
optimistic or pessimistic bias.
·
The test set doesn’t
effect the outcome
of learning process
which only uses training set.
Validation Set:
· Validation set is not used for learning but is for avoiding overfitting.
· The idea of ‘Validation Set’ is identical to ‘Test Set’ but there is a difference.
· We remove a subset from data and this subset is not used for training.
· Although validation set is not directly used for training, it will be used in making certain choices in the learning process. As the set affects the learning process, it is no longer a test set.
· The more we use validation set to fine tune the model, the more validation set becomes like a training set.
· The validation and test sets are roughly the same sizes, much smaller than the size of the training set. The learning algorithm cannot use examples from these two subsets to build the model.
· Validation set is used for selecting the model complexity.
In the past, the rule of thumb was to use 70% of the dataset for training, 15% for validation and 15% for testing. However, in the age of big data, datasets often have millions of examples. In such cases, it could be reasonable to keep 95% for training and 2.5%/2.5% for validation/testing.
3. MACHINE LEARNING FRAMEWORK:
· y = f(x)
o
y = output, x = feature
representation, f( ) is prediction function
· Training: Given a training set, estimate the prediction function f( )
by minimizing the prediction error.
· Testing: Apply f ( ) to unknown
test sample x and predicted
value (output) is y.
·
y = f(w, x) (for linear model)
o
y = output, x = feature representation, w = weight,
f(w, ) = prediction function
· Training: Given a training set, estimate
the prediction function, f( ) by minimizing the prediction error.
· Testing: Apply f(w, ) to unknown test sample x and predicted
value (output) is y
o
Parameter: primary
problem is to find the parameters W.
AI/ML Problem – phases:
The following are the various
steps involved in problem solving
using AI/ML:
·
Define your task
· Collect Data
·
Preprocessing of Data
·
Dimensionality Reduction/Feature Selection
· Choose ML Algorithm
· Experimental Design
· Test and Validate
·
Run System
MACHINE LEARNING
WORK FLOW:
2. MACHINE LEARNING - INTRODUCTION
DEFINITION:
Machine
Learning is a set of methods that can automatically detect patterns in data and then use the uncovered patterns
to predict future data or to perform other kinds of decision making under uncertainty (such as planning
how to collect more data).
LEARNING PROBLEM:
A computer program is said
to learn from experience E with respect to some
class of tasks T and performance measure P, if its
performance at tasks in T,
as measured by P, improves with experience E.
Example: a computer program
that learns to play checkers P in this case is measured by its ability to win
T in this case is playing checkers games
E in this case is obtained by playing games against itself
In general,
to have a well-defined learning problem we must
identify these three features:
1.
The class of tasks (T)
2.
The measure of performance to be improved
(P)
3.
The source of experience (E)
Checkers learning
problem:
·
Task T – playing
checkers
·
Performance measure
P – percent of games
won against opponents
· Training experience E – playing
practice games against
itself
Handwriting recognition learning
problem:
·
Task T – recognizing and classifying
handwritten words within
images
· Performance measure P – percentage of words correctly classified
· Training Experience E – a database of handwritten words with given classifications
Robot driving learning problem:
·
Task T – driving on
public four lane highways using vision
sensors
· Performance measure P – average distance travelled before an
error (as judged by human
overseer)
· Training experience E – a sequence of images and steering commands
recorded while observing a human
driver
E-mail classification:
·
Task T: Categorize email messages as spam or legitimate.
· Performance P: Percentage of email messages correctly classified.
· Training Experience E: Database of emails, some with human-given labels
1.Machine Learning – TABLE OF CONTENTS
1. Machine Learning – Introduction
2.
Machine Learning Framework
3.
AI/ML Problem – phases
4.
ML Workflow, Tools and Landscape
5.
Dataset and subsets
6.
Machine Learning Types
7.
Supervised Learning
8.
Supervised Learning
Types
9.
Classification Vs
Regression
10. Performance metrics
11. Classification Algorithms
a. Logistic Regression
b. K-Nearest Neighbor (KNN)
c. Decision Tree
d. Random Forest – Gradient Boosting
e. NaĆÆve Bayes
f. Support Vector Machines
(SVM)
g. Gradient Descent
h. Neural Networks
12. Neural Networks
a. Perceptron
b. Multi-Level Perceptron (MLP)
c. Artificial Neural Networks
(ANN)
d. Deep Neural Networks
(DNN)
13. Activation Functions
14. Dropout
15. Backpropagation
16. Cross Validation
17. Model Compression
18. Loss Function
19. Deep Learning/Deep Neural Networks
a. Convolutional Neural networks
(CNN)
b. Recurrent Neural Networks
(RNN)
c. Restricted Boltzmann Machine
(RBM)
d. Deep Belief Networks
(DBN)
e. Autoencoders
20. Convolutional Neural Networks
a. Core idea
b. Convolution layer
c. Various terms associated with CNN
d. Pooling layer and types of pooling
e. CNN Architecture
f. Optimization of CNN
g. Working of CNN
h. Applications
21.
Recurrent Neural
Networks
a.
RNN and limitations
b. Differences with CNN
c. Long Short-Term Memory
(LSTM)
d. Gated Feedback Recurrent
Neural Networks (GRU)
22. Restricted Boltzmann Machine
(RBM)
23. Deep Belief Networks
(DBN)
24. Autoencoders
a. Types
b. Applications
c. Convolutional Autoencoders
d. Denoising Autoencoders
e. Deep Autoencoders
f. Variational Autoencoders (VAE)
g. Sparse Autoencoders
25. Siamese Neural Networks
26. Generative Adversarial Networks
(GAN)
27. Regression Algorithms
a. Linear Regression
b. Multiple Linear Regression
c. Polynomial Regression
d. Ridge Regression
e. Lasso Regression
f. ElasticNet Regression
28. Unsupervised Learning
29. Unsupervised Learning Algorithms
a. Clustering
b. K-Means Clustering
c. Hierarchical Clustering
d. DBSCAN
e. Gaussian Mixtures
f. Spectral Clustering
30. Dimensionality Reduction
a. Feature Engineering
b. Multicollinearity
c. Factor Analysis
d. Principal Component Analysis
(PCA)
e. Linear Discriminant Analysis
(LDA)
f. Isometric Mapping (IsoMap)
g. Locally Linear Embedding (LLE)
h. t-distribution Stochastic Neighborhood Embedding (t-SNE)
31. Semi-supervised Learning
32. Semi-supervised learning algorithms:
a. Self Training
b. Generative Models
c. S3VMs
d. Graph Based Algorithms
e. Multiview algorithms
33. Reinforcement Learning
a. Introduction
b. Approaches
c. Markov Decision Process
d. Q-learning
e. Application of RL
34. Other topics
a. Word2Vec
b. Generalization
c. Overfitting
d. Regularization
e. Bias
f. Variance
g. Occam’s Razor
h. Bag of Words (BoW)
i. Recommender Systems
j. Ensemble Methods
k. Natural Language Processing (NLP)