<< Chapter < Page Chapter >> Page >

Numerical identification of outliers: calculating s And finding outliers manually

If you do not have the function LinRegTTest, then you can calculate the outlier in the first example by doing the following.

First, square each | y ŷ |

The squares are

  • 35 2
  • 17 2
  • 16 2
  • 6 2
  • 19 2
  • 9 2
  • 3 2
  • 1 2
  • 10 2
  • 9 2
  • 1 2

Then, add (sum) all the | y ŷ | squared terms using the formula

Σ i   =   1 11 ( | y i y ^ i | ) 2 = Σ i   =   1 11 ε i 2 (Recall that y i ŷ i = ε i .)

= 35 2 + 17 2 + 16 2 + 6 2 + 19 2 + 9 2 + 3 2 + 1 2 + 10 2 + 9 2 + 1 2

= 2440 = SSE . The result, SSE is the Sum of Squared Errors.

Next, calculate s , the standard deviation of all the y ŷ = ε values where n = the total number of data points.

The calculation is s = SSE n 2 .

For the third exam/final exam problem, s = 2440 11 2 = 16.47 .

Next, multiply s by 2:
(2)(16.47) = 32.94
32.94 is 2 standard deviations away from the mean of the y ŷ values.

If we were to measure the vertical distance from any data point to the corresponding point on the line of best fit and that distance is at least 2 s , then we would consider the data point to be "too far" from the line of best fit. We call that point a potential outlier .

For the example, if any of the | y ŷ | values are at least 32.94, the corresponding ( x , y ) data point is a potential outlier.

For the third exam/final exam problem, all the | y ŷ |'s are less than 31.29 except for the first one which is 35.

35>31.29 That is, | y ŷ | ≥ (2)(s)

The point which corresponds to | y ŷ | = 35 is (65, 175). Therefore, the data point (65,175) is a potential outlier. For this example, we will delete it. (Remember, we do not always delete an outlier.)

Note

When outliers are deleted, the researcher should either record that data was deleted, and why, or the researcher should provide results both with and without the deleted data. If data is erroneous and the correct values are known (e.g., student one actually scored a 70 instead of a 65), then this correction can be made to the data.



The next step is to compute a new best-fit line using the ten remaining points. The new line of best fit and the correlation coefficient are:

ŷ = –355.19 + 7.39 x and r = 0.9121

Using this new line of best fit (based on the remaining ten data points in the third exam/final exam example ), what would a student who receives a 73 on the third exam expect to receive on the final exam? Is this the same as the prediction made using the original line?

Using the new line of best fit, ŷ = –355.19 + 7.39(73) = 184.28. A student who scored 73 points on the third exam would expect to earn 184 points on the final exam.

The original line predicted ŷ = –173.51 + 4.83(73) = 179.08 so the prediction using the new line with the outlier eliminated differs from the original prediction.

Got questions? Get instant answers now!
Got questions? Get instant answers now!

Try it

The data points for the graph from the third exam/final exam example are as follows: (1, 5), (2, 7), (2, 6), (3, 9), (4, 12), (4, 13), (5, 18), (6, 19), (7, 12), and (7, 21). Remove the outlier and recalculate the line of best fit. Find the value of ŷ when x = 10.

ŷ = 1.04 + 2.96 x ; 30.64

Got questions? Get instant answers now!

The Consumer Price Index (CPI) measures the average change over time in the prices paid by urban consumers for consumer goods and services. The CPI affects nearly all Americans because of the many ways it is used. One of its biggest uses is as a measure of inflation. By providing information about price changes in the Nation's economy to government, business, and labor, the CPI helps them to make economic decisions. The President, Congress, and the Federal Reserve Board use the CPI's trends to formulate monetary and fiscal policies. In the following table, x is the year and y is the CPI.

Questions & Answers

if three forces F1.f2 .f3 act at a point on a Cartesian plane in the daigram .....so if the question says write down the x and y components ..... I really don't understand
Syamthanda Reply
hey , can you please explain oxidation reaction & redox ?
Boitumelo Reply
hey , can you please explain oxidation reaction and redox ?
Boitumelo
for grade 12 or grade 11?
Sibulele
the value of V1 and V2
Tumelo Reply
advantages of electrons in a circuit
Rethabile Reply
we're do you find electromagnetism past papers
Ntombifuthi
what a normal force
Tholulwazi Reply
it is the force or component of the force that the surface exert on an object incontact with it and which acts perpendicular to the surface
Sihle
what is physics?
Petrus Reply
what is the half reaction of Potassium and chlorine
Anna Reply
how to calculate coefficient of static friction
Lisa Reply
how to calculate static friction
Lisa
How to calculate a current
Tumelo
how to calculate the magnitude of horizontal component of the applied force
Mogano
How to calculate force
Monambi
a structure of a thermocouple used to measure inner temperature
Anna Reply
a fixed gas of a mass is held at standard pressure temperature of 15 degrees Celsius .Calculate the temperature of the gas in Celsius if the pressure is changed to 2×10 to the power 4
Amahle Reply
How is energy being used in bonding?
Raymond Reply
what is acceleration
Syamthanda Reply
a rate of change in velocity of an object whith respect to time
Khuthadzo
how can we find the moment of torque of a circular object
Kidist
Acceleration is a rate of change in velocity.
Justice
t =r×f
Khuthadzo
how to calculate tension by substitution
Precious Reply
hi
Shongi
hi
Leago
use fnet method. how many obects are being calculated ?
Khuthadzo
khuthadzo hii
Hulisani
how to calculate acceleration and tension force
Lungile Reply
you use Fnet equals ma , newtoms second law formula
Masego
please help me with vectors in two dimensions
Mulaudzi Reply
how to calculate normal force
Mulaudzi
Got questions? Join the online conversation and get instant answers!
Jobilize.com Reply

Get Jobilize Job Search Mobile App in your pocket Now!

Get it on Google Play Download on the App Store Now




Source:  OpenStax, Introductory statistics. OpenStax CNX. May 06, 2016 Download for free at http://legacy.cnx.org/content/col11562/1.18
Google Play and the Google Play logo are trademarks of Google Inc.

Notification Switch

Would you like to follow the 'Introductory statistics' conversation and receive update notifications?

Ask