# 2.1 Machine learning lecture 2 course notes  (Page 3/6)

 Page 3 / 6
$\begin{array}{ccc}\hfill \Phi & =& \frac{1}{m}\sum _{i=1}^{m}1\left\{{y}^{\left(i\right)}=1\right\}\hfill \\ \hfill {\mu }_{0}& =& \frac{{\sum }_{i=1}^{m}1\left\{{y}^{\left(i\right)}=0\right\}{x}^{\left(i\right)}}{{\sum }_{i=1}^{m}1\left\{{y}^{\left(i\right)}=0\right\}}\hfill \\ \hfill {\mu }_{1}& =& \frac{{\sum }_{i=1}^{m}1\left\{{y}^{\left(i\right)}=1\right\}{x}^{\left(i\right)}}{{\sum }_{i=1}^{m}1\left\{{y}^{\left(i\right)}=1\right\}}\hfill \\ \hfill \Sigma & =& \frac{1}{m}\sum _{i=1}^{m}\left({x}^{\left(i\right)}-{\mu }_{{y}^{\left(i\right)}}\right){\left({x}^{\left(i\right)}-{\mu }_{{y}^{\left(i\right)}}\right)}^{T}.\hfill \end{array}$

Pictorially, what the algorithm is doing can be seen in as follows:

Shown in the figure are the training set, as well as the contours of the two Gaussian distributions that have been fit to the data in each of thetwo classes. Note that the two Gaussians have contours that are the same shape and orientation, since they share a covariance matrix $\Sigma$ , but they have different means ${\mu }_{0}$ and ${\mu }_{1}$ . Also shown in the figure is the straight line giving the decision boundary at which $p\left(y=1|x\right)=0.5$ . On one side of the boundary, we'll predict $y=1$ to be the most likely outcome, and on the other side, we'll predict $y=0$ .

## Discussion: gda and logistic regression

The GDA model has an interesting relationship to logistic regression. If we view the quantity $p\left(y=1|x;\Phi ,{\mu }_{0},{\mu }_{1},\Sigma \right)$ as a function of $x$ , we'll find that it can be expressed in the form

$p\left(y=1|x;\Phi ,\Sigma ,{\mu }_{0},{\mu }_{1}\right)=\frac{1}{1+exp\left(-{\theta }^{T}x\right)},$

where $\theta$ is some appropriate function of $\Phi ,\Sigma ,{\mu }_{0},{\mu }_{1}$ . This uses the convention of redefining the ${x}^{\left(i\right)}$ 's on the right-hand-side to be $n+1$ -dimensional vectors by adding the extra coordinate ${x}_{0}^{\left(i\right)}=1$ ; see problem set 1. This is exactly the form that logistic regression—a discriminative algorithm—used to model $p\left(y=1|x\right)$ .

When would we prefer one model over another? GDA and logistic regression will, in general, give different decision boundaries when trained on the same dataset. Which is better?

We just argued that if $p\left(x|y\right)$ is multivariate gaussian (with shared $\Sigma$ ), then $p\left(y|x\right)$ necessarily follows a logistic function. The converse, however, is not true; i.e., $p\left(y|x\right)$ being a logistic function does not imply $p\left(x|y\right)$ is multivariate gaussian. This shows that GDA makes stronger modeling assumptions about the data than does logistic regression. It turns out that when these modelingassumptions are correct, then GDA will find better fits to the data, and is a better model. Specifically, when $p\left(x|y\right)$ is indeed gaussian (with shared $\Sigma$ ), then GDA is asymptotically efficient . Informally, this means that in the limit of very large training sets (large $m$ ), there is no algorithm that is strictly better than GDA (in terms of, say, how accurately they estimate $p\left(y|x\right)$ ). In particular, it can be shown that in this setting, GDA will be a better algorithm than logistic regression; and more generally,even for small training set sizes, we would generally expect GDA to better.

In contrast, by making significantly weaker assumptions, logistic regression is also more robust and less sensitive to incorrect modeling assumptions. There are many different sets of assumptions that would lead to $p\left(y|x\right)$ taking the form of a logistic function. For example, if $x|y=0\sim \mathrm{Poisson}\left({\lambda }_{0}\right)$ , and $x|y=1\sim \mathrm{Poisson}\left({\lambda }_{1}\right)$ , then $p\left(y|x\right)$ will be logistic. Logistic regression will also work well on Poisson data like this. But if we were to use GDA on such data—and fit Gaussian distributions tosuch non-Gaussian data—then the results will be less predictable, and GDA may (or may not) do well.

To summarize: GDA makes stronger modeling assumptions, and is more data efficient (i.e., requires less training data to learn “well”)when the modeling assumptions are correct or at least approximately correct. Logistic regression makes weaker assumptions, and is significantly more robust to deviationsfrom modeling assumptions. Specifically, when the data is indeed non-Gaussian, then in the limit of large datasets, logistic regression will almost always do better thanGDA. For this reason, in practice logistic regression is used more often than GDA. (Some related considerations about discriminative vs. generative models also apply forthe Naive Bayes algorithm that we discuss next, but the Naive Bayes algorithm is still considered a very good, and is certainly also a very popular, classification algorithm.)

what is variations in raman spectra for nanomaterials
I only see partial conversation and what's the question here!
what about nanotechnology for water purification
please someone correct me if I'm wrong but I think one can use nanoparticles, specially silver nanoparticles for water treatment.
Damian
yes that's correct
Professor
I think
Professor
what is the stm
is there industrial application of fullrenes. What is the method to prepare fullrene on large scale.?
Rafiq
industrial application...? mmm I think on the medical side as drug carrier, but you should go deeper on your research, I may be wrong
Damian
How we are making nano material?
what is a peer
What is meant by 'nano scale'?
What is STMs full form?
LITNING
scanning tunneling microscope
Sahil
how nano science is used for hydrophobicity
Santosh
Do u think that Graphene and Fullrene fiber can be used to make Air Plane body structure the lightest and strongest. Rafiq
Rafiq
what is differents between GO and RGO?
Mahi
what is simplest way to understand the applications of nano robots used to detect the cancer affected cell of human body.? How this robot is carried to required site of body cell.? what will be the carrier material and how can be detected that correct delivery of drug is done Rafiq
Rafiq
what is Nano technology ?
write examples of Nano molecule?
Bob
The nanotechnology is as new science, to scale nanometric
brayan
nanotechnology is the study, desing, synthesis, manipulation and application of materials and functional systems through control of matter at nanoscale
Damian
Is there any normative that regulates the use of silver nanoparticles?
what king of growth are you checking .?
Renato
What fields keep nano created devices from performing or assimulating ? Magnetic fields ? Are do they assimilate ?
why we need to study biomolecules, molecular biology in nanotechnology?
?
Kyle
yes I'm doing my masters in nanotechnology, we are being studying all these domains as well..
why?
what school?
Kyle
biomolecules are e building blocks of every organics and inorganic materials.
Joe
anyone know any internet site where one can find nanotechnology papers?
research.net
kanaga
sciencedirect big data base
Ernesto
Introduction about quantum dots in nanotechnology
what does nano mean?
nano basically means 10^(-9). nanometer is a unit to measure length.
Bharti
do you think it's worthwhile in the long term to study the effects and possibilities of nanotechnology on viral treatment?
absolutely yes
Daniel
how did you get the value of 2000N.What calculations are needed to arrive at it
Privacy Information Security Software Version 1.1a
Good
Got questions? Join the online conversation and get instant answers!    By      