Mathematics Study Guide for the TABE Test
Page 18
Measurement, Data, and Probability: Statistics and Probability
A statistical question will produce multiple and varied answers. So, if you ask Marie what she scored on a math test, that’s not a statistical question. There’s only one answer to that question. However, if you ask how Marie’s whole class scored on the test, that is a statistical question, because every student will provide a different answer. You could then combine their answers into one number or grade to represent the whole class.
That is the role of statistics: collecting a range of data to represent a whole.
Samples
When researchers are collecting data about a certain group, they have to decide how many members of the group to include in their study. If, for instance, you wanted the average height of a high school senior in Ohio, you wouldn’t be able to measure every high school senior in the state, so you’d measure a representative group called a sample. Maybe you would measure \(500\) seniors. The sample size would be \(500\). The bigger the sample size the more accurate your results will be.
Random Sample
It’s important for the accuracy of a study for the sample to be randomly selected. For example, say you were doing a study of how many people in the general population were sick. You wouldn’t want to pick your sample from a medical clinic waiting room, because those people are more likely to be sick than someone out in the general public. That would make the number of people with an illness higher than in the average population. Along with larger sample sizes, randomization ensures better research results.
Multiple Samples and Variability
Another way to increase the reliability of the statistics is measuring multiple (or simulated) samples. Having more than one sample allows us to make comparisons and learn more information.
The table below shows the result of a study in which the vowels in randomly selected words were counted. There were three distinct samples, each of which consisted of \(30\) words:
| a | e | i | o | u | |
|---|---|---|---|---|---|
| Sample \(1\) | \(7\) | \(14\) | \(3\) | \(6\) | \(4\) |
| Sample \(2\) | \(10\) | \(12\) | \(8\) | \(3\) | \(3\) |
| Sample \(3\) | \(4\) | \(12\) | \(5\) | \(6\) | \(7\) |
| Average | \(7\) | \(12.7\) | \(5.3\) | \(5\) | \(4.7\) |
It’s pretty clear by looking at the averages that the most used vowel was e, with a coming in second and u coming in last.
There’s more valuable information that can be gleaned from the table. For one, we can study the variability (or spread) of the data, meaning how much the numbers for each vowel stray from the average. Look at the e column data and you will see the numbers \(14, 12,\) and \(12\). All three are quite close to the average of \(12.7\). In other words, there’s not much variability among them.
On the other hand, look at the a column data: \(7, 10,\) and \(4\), with an average of \(7\). The \(7\) is right on the average, but the \(4\) and \(10\) aren’t that close. In other words, the data in that column have greater variability than the e column data.
In statistics, low variability is good. If all the data points are very close to each other, you can be pretty confident about your statistics. High variability is bad. If the numbers are all over the place, the data likely isn’t very reliable.
Inference
An inference is a conclusion or prediction about an entire population based on information collected from a sample. Because it is often impossible or impractical to gather data from every member of a population, researchers use representative samples to estimate what is true about the larger group.
For example, suppose a researcher randomly surveys \(500\) high school seniors in a state and finds that \(70\%\) plan to attend college after graduation. The researcher may then infer that approximately \(70\%\) of all high school seniors in the state plan to attend college. This conclusion is an inference because it is based on a sample rather than the entire population.
The accuracy of an inference depends heavily on how the sample was chosen. A representative sample reflects the characteristics of the population being studied. Random sampling helps achieve this because every member of the population has an equal chance of being selected. When samples are representative, the conclusions drawn from them are more likely to be valid.
On the other hand, if a sample is biased, the resulting inference may be misleading. For example, if a survey about student study habits is conducted only among students enrolled in advanced classes, the results may not accurately represent all students at the school. This is because the sample does not reflect the entire population.
When evaluating statistical claims, it is important to ask whether the sample was large enough, randomly selected, and representative of the population. These factors help determine whether an inference is likely to be reliable.
Let’s look at an example problem involving inference.
A school wants to estimate the average number of hours students spend on homework each week. A random sample of \(100\) students reports an average of six hours per week. What inference can the school make?
Solution
Because the students were selected randomly, the school can infer that the average homework time for the entire student population is approximately six hours per week.
However, this conclusion is still an estimate. Different random samples might produce slightly different results, which is why larger and more representative samples generally lead to more reliable inferences.
Measures of Center
When answering a statistical question, there is always the question of how to show the data in a way that will be useful. You’ve learned about graphing already, but there are additional concepts that come into play when graphing statistics.
Mean, Median, and Mode
Sometimes it’s nice to be able to boil a whole data set down to one number to represent the whole set. This is useful when you want to compare different data sets. Suppose you weigh \(100\) horses and \(100\) cows to determine if a typical cow weighs more or less than a typical horse. Rather than trying to figure this out by comparing all \(200\) weights, it would be useful to find a single representative value for the whole. Luckily, we have a few such values to choose from.
Often in statistics, you will find a central tendency (or central value) to summarize a large data set. There are three common central values: mean, median, and mode.
The mean (or average) of a data set is the result of dividing the sum of all the values by the total number of values. So, suppose you are given this data set:
\[40, \,56, \,38,\, 67,\, 50, \,40, \,55, \, 70\]Simply add up all the values and divide by \(8\) to get its mean:
\[\frac{40 + 56 + 38 + 67 + 50 + 40 + 55 + 70}{8} = \frac {416}{8} = 52\]The median of a data set is the middle value when all the values are arranged in ascending or descending order. When the data set has an odd number of values, there will be one value in the middle. When the data set has an even number of values, there will be two values in the middle, and you will need to find their average (by adding them and dividing by \(2\)). Let’s find the median of the list from above by first sorting it into ascending order:
\[38, \,40,\, 40,\, 50, \,55, \,56, \,67, \,70\]We know there are eight values, so we’ll have two values in the middle:
\[50, 55\]Therefore, the median of this data set is:
\[(50 + 55) \div 2 = 52.5\]The mode of a data set is the value that occurs most frequently. In the set we’ve been using, \(40\) occurs twice while the other values only occur once each. Therefore, the mode is \(40\). You can have two modes (bimodal), three modes (trimodal), or more. You can also have no mode if no single value is more common than any other.
Distribution
Suppose you measure the nose length of \(1\text{,}000\) randomly chosen people. Some people will have longer noses and some will have shorter ones. A few will have very long noses and a few will have very short noses. Assuming you’ve selected a sufficient sample size, you can expect most of the nose lengths to be near the average of all nose lengths.
If you graph the numbers of people versus their nose lengths, you should expect to get a graph of the distribution of the values that looks a lot like the one below:

This is called a bell curve. The dashed line in the center represents the average nose length. The fact that the graph is so high there tells us that more people have that nose length than any other length. The shape of the graph shows that nose lengths farther from the center are less common. The graph being very low on both ends shows that very few people have really long noses or really short noses.
In reality, the graph probably wouldn’t be perfectly symmetrical like this one, but it illustrates how the distribution of data generally works. The closer you get to the average, the more values you’ll find.
Central Tendency vs. Variability
Measures of center and measures of variation describe different features of a data set.
A measure of center summarizes an entire set of data with a single value that represents a typical data point. The most common measures of center, as you have learned, are the mean, median, and mode. These values help us identify where the data are generally located.
A measure of variation (or spread) describes how much the data values differ from one another. Two data sets can have the same mean but be very different in how spread out they are.
Consider these two data sets:
- Set A: \(48,\ 49,\ 50,\ 51,\ 52\)
- Set B: \(10,\ 30,\ 50,\ 70,\ 90\)

Both sets have a mean of \(50\). However, Set A is tightly clustered around the center, while Set B is spread over a much wider range of values. The measure of center is the same, but the variation is very different.
One simple measure of variation is the range, which is found by subtracting the smallest value from the largest value.
For Set A, this is the range:
\[52-48=4\]For Set B, this is the range:
\[90-10=80\]The much larger range of Set B tells us that its values vary a great deal more.
The shape of a distribution can also provide information about variation. A perfectly balanced bell-shaped distribution is called symmetric because the left and right sides mirror each other.

However, real-world data sets are typically not perfectly symmetric. When one side of a distribution stretches farther than the other, the distribution is said to be skewed:
-
A right-skewed distribution has a longer tail on the right side. This often occurs when a few unusually large values pull the mean upward.
-
A left-skewed distribution has a longer tail on the left side. This occurs when a few unusually small values pull the mean downward.

For example, income data are often right-skewed because a small number of people earn much more than the majority of the population.
When analyzing data, it is important to consider both the center and the variability. The center tells us what is typical, while the variability tells us how consistent or spread out the data are. Together, they provide a more complete picture of the distribution.
Box Plot
Box plots, also known as box and whisker plots, are a special type of data display that rely on measures of central tendency. They have a box in the middle and lines (“the whiskers”) extending out on either end. This kind of graph lets you see if there is a pattern in the data. Box plots are made up of four quarters and look like this:

Retrieved from: https://openstax.org/books/statistics/pages/2-4-box-plots, Figure 2.13
The box plot above is made up of the shoe sizes of the customers in a shoe store over a single day, which were:
\[1, \,1,\, 2, \,2,\, 4, \,6, \,7,\, 7,\, 8,\, 8.5,\, 9, \,10, \,10, \,11.5\]The whisker on the left represents one quarter of all the sizes and shows that one quarter of all sizes that day were from \(1\) to \(2\). The dotted line that divides the box represents another quarter, meaning one quarter of the shoes sold that day were from size \(2\) to \(7\). The box section to the right of the dotted line is another quarter of the sales, those from size \(7\) to \(9\). The whisker to the right is the fourth and final quarter, meaning one-fourth of the shoes sold on that day were from size \(9\) to \(11.5\).
Outliers
Imagine seeing this graph below:

What’s up with dog Ten? Its weight is so far away from the other weights that something weird must have happened. Maybe the scale malfunctioned. Maybe that dog was just a puppy. This kind of weird result is called an outlier. Whatever the reason for it, we probably can’t trust it much. It increases the variability (or spreads the data out more). Most likely we would just throw it out of our data set. If we left it in, it would drag the average down below what it likely should be.
Clustering
When there is an enormous population to study, researchers sometimes group this large number into smaller groups and sample those groups during a study. These groups are called clusters.
For example, suppose a researcher wants to study student opinions across an entire state. Surveying every student would be difficult and expensive. Instead, the researcher might divide the state into school districts and randomly select a few districts to survey. Each selected district would be a cluster.

Cluster sampling can save time and money because researchers can collect data from groups rather than from individuals scattered across a large area.
The word clustering can also describe a pattern in data. When data points tend to gather in certain regions of a graph instead of being spread evenly, they are said to form clusters. For example, suppose a teacher records the number of hours students studied for an exam. Suppose the following study hours were reported by the students:
\[1,\ 2,\ 2,\ 3,\ 3,\ 3,\ 4,\ 4,\ 4,\ 4,\ 7,\ 8\]You can see that most of the values are grouped between two and four hours, forming a cluster. The values seven and eight lie outside the cluster and may represent unusual cases.
Recognizing clusters helps researchers identify patterns within data. Clusters can reveal common behaviors, preferences, or characteristics within a population.
Positive and Negative Associations
An association in statistics is a relationship between two variables. Think about skin wrinkles and how a person gets more as they get older. The association between age and wrinkles is called a positive association because they both change in the same direction. As a general rule, as one’s age increases, so do their wrinkles.
On the other hand, a negative association, as you might guess, is the opposite. For example, think of your car’s odometer and gas gauge as you drive. As your miles increase, the gas in your tank decreases.
We use these types of associations in statistics all the time to describe how a change in one variable relates to a change in another. Remember the graph we saw earlier that displayed the scores students got on two tests? It included a best-fit line:

Retrieved from: https://openstax.org/books/statistics/pages/12-2-the-regression-equation Figure 12.7
We found that there was a trend: The higher a student scored on one test was predictive of a higher score on the other test. Another way of saying that is there is a positive association between the scores on both tests.
Correlation and Causation
Another word for association is correlation. The variables being studied change in ways that are connected (correlated) to each other. The correlation can be either positive, negative, or nonexistent.
A correlation coefficient tells us the degree and type of correlation on a scale of \(-1.0\) to \(1.0\). These are the types of correlations possible:
-
At \(0.0\), there is no correlation at all.
-
Greater than \(0.0\) to \(1.0\) indicates a positive correlation, which is stronger as the value rises.
-
Less than \(0.0\) and down to \(-1.0\) indicates a negative correlation, which is stronger as the value decreases.
Causation means that a change in one quantity causes a change in the other. Driving a greater distance will cause your car to use more gas. Therefore, we can say that driving causes the gas gauge to go down. That’s a simple cause-and-effect example from the real world.
However, correlation does not always mean causation. In other words, even when two variables are correlated (or associated), it does not mean one variable caused the other variable. There could be other variables involved that aren’t being measured or studied.
A well-known example of this principle is the correlation between ice cream sales and the murder rate in the summer months. Both go up when it’s hotter. If you were to graph ice cream sales against the murder rate, you would see a very strong positive correlation. They both go up. Does that mean eating ice cream causes murders? Of course not. Even though there is a strong correlation, there is no causation.
Linear and Nonlinear Association
One way to show associations is by using a graph. When we put data on a graph, if there is an association, the graph will either be a straight line or have a curve. A graph of a straight line is called a linear association (you learned about linear equations earlier). By contrast, a nonlinear association will not be a straight line, but rather it will curve.
An example of a linear association is distance and time when you’re moving at a constant speed. Every mile will take the same amount of time. On the other hand, if you drop something, it actually picks up speed as it falls. Each foot that it falls takes a little less time than the one before it. Illustrating the change in speed over time will result in a curved graph:
Probability
Probability is a mathematical concept that relates to predicting or calculating the chance that an event will occur. It is often given by a decimal number between \(0\) and \(1\), where \(0\) means the event won’t happen and \(1\) means it definitely will happen. It can also be given as a ratio, such as \(\frac{3}{5}\), or three out of five.
Probability is calculated by dividing the number of desired (or favored) outcomes (\(n(A)\)) by the total number of possible outcomes (\(n(S)\)), which can be reduced to the following formula:
\[P(E) = \frac{n(A)}{n(S)}\]Suppose a clown has prepared \(10\) tricks for a child’s birthday party, including pulling a bunny out of a hat. If the time he has for his performance only allows him to perform one trick, what is the likelihood that he will perform the bunny trick?
Calculating probability requires simple division, but first we have to determine which number is being divided by which other number. In this case, our desired outcome is one single event, the bunny trick, and the total of possible outcomes is every trick the clown has prepared, which is \(10\). So, the probability is:
\[\frac{1}{10}=0.1\]What if the clown had time to do two tricks? Then we simply multiply our original probability by \(2\) to get \(0.2\).
Note: The term “desired” in this context doesn’t necessarily mean wanted, but rather intended.
Probability Model
Using a model to visualize a probability problem helps make it more concrete and easier to understand. You will need to be able to come up with a reasonable idea of what the probability is of a certain simple event happening. For example, suppose there are three red marbles, four green marbles, and nine blue marbles in a cloth bag. If you reach in and randomly pull a marble out, what is the probability that it will be green?
Recall that probability is:
\[\frac {\text{number of favorable outcomes}} {\text{total number of possible outcomes}}\]The bag has \(16\) marbles in it, so there are only \(16\) outcomes possible when you pull one out. Green is the favorable outcome and of the \(16\) possible outcomes, there are \(4\) that will be green. That gives us a reasonable guess for the probability:
\[\frac{4}{16} = \frac{1}{4} = 0.25\]If we were to actually try this with marbles, would we get one green marble every time we picked four marbles? If not, can you think of any reason why not?
The reality is that probability only tells us what is most likely to happen. It tells us that in the very long run we will get really close to getting a green for every four marbles we pick. In the short run, it’s unpredictable. There might be something weird happening, like maybe green marbles are lighter than the other ones so they tend to rise to the top. Far fetched? Maybe, but try to think of things like this to try to explain why real observations don’t agree with a calculated probability.
Relative Frequency
While theoretical probability is calculated using a formula, relative frequency is based on actual results from an experiment or observation. Relative frequency tells us how often an event occurs compared to the total number of trials.
We calculate relative frequency by taking the total number of times an event occurred and dividing it by the total number of trials. You can think of it as this formula:
\[F = \frac{E}{T}\]As the number of trials increases, the relative frequency usually gets closer and closer to the theoretical probability. This idea is known as the long-run relative frequency of an event.
For example, suppose a fair coin is flipped \(100\) times and lands on heads \(47\) times. The relative frequency of heads is:
\[\frac{47}{100}=0.47\]Since a coin has only two sides, the theoretical probability of getting heads on a fair coin is:
\[\frac{1}{2}=0.50\]Although \(0.47\) is not exactly \(0.50\), it is fairly close. If the coin were flipped thousands of times, the relative frequency would likely move even closer to \(0.50\).
Predicting Relative Frequency
If you know the probability of an event and the number of times an experiment will be performed, you can predict approximately how often the event should occur. Look at the example below.
A spinner is divided into four equal sections, one of which is red. If the spinner is spun \(200\) times, approximately how many times should it land on red?
Solution
The probability of landing on red is:
\[P_r=\frac{1}{4}=0.25\]Multiply the probability by the number of spins:
\[0.25 \times 200 = 50\]We would expect the spinner to land on red approximately \(50\) times.
Conditional Relative Frequency
A conditional relative frequency is the relative frequency of one category within a specific group. In a two-way table, it is found by dividing a value in the table by the total for the row or column that defines that group.
Words such as among or of those who can help identify the group being considered. The total for that group becomes the denominator.
For example, suppose \(20\) students attend an after-school program, and \(14\) of those students also participate in a school club. Can you determine the conditional relative frequency of participating in a school club among students who attend the after-school program?
Because the question asks about students who attend the after-school program, \(20\) is the total for the group being considered. Of those students, \(14\) participate in a school club:
\[\frac{14}{20}=0.70=70\%\]Therefore, the conditional relative frequency is \(70\%\).
Note: For a conditional relative frequency, use the total for the specified group as the denominator, not the total for the entire data set.
Comparing a Probability Model to Observed Results
A probability model predicts what should happen, while observed frequencies tell us what actually happened. Comparing the two helps us determine whether the model is a good representation of reality. Let’s try an example.
A bag contains five red marbles and five blue marbles. The probability of drawing a red marble is \(\frac{5}{10}=\frac{1}{2}\). A student replaces the marble after each draw and performs \(40\) draws. The student records \(14\) red marbles and \(26\) blue marbles. How do the results compare to the model?
Solution
The probability model predicts \(\frac{1}{2} \times 40 = 20\) red marbles, for a predicted probability of \(0.5\).
The student observed only \(14\) red marbles.
The observed relative frequency is:
\[\frac{14}{40}=0.35\]This is lower than the predicted probability of \(0.5\). There are several factors that can cause discrepancies between expected and observed results:
- The number of trials may be too small.
- Random chance may produce unusual short-term results.
- The experiment may not have been conducted fairly.
- The probability model may not accurately represent the situation.
When many more trials are performed, observed frequencies generally move closer to the probabilities predicted by the model. If they don’t, then the likelihood is that the model is wrong.
The key idea is that probability predicts what should happen in the long run, while relative frequency measures what actually happened. The more trials that are conducted, the more closely the relative frequency tends to match the theoretical probability.
Combinations and Permutations
Combination in probability refers to a method for counting the possible number of ways a thing can be done or the pairing of things where the order of the elements doesn’t matter. Combinations can be calculated by using this formula:
\[C = \frac {n!}{r!(n-r)!}\]where \(n\) is the number of things to choose from and \(r\) is the number of choices. The exclamation point (\(!\)) is the factorial symbol, which tells you to multiply every positive integer from \(1\) up to the number in front of the symbol. So, if you see \(4!\), that means:
\[4\cdot 3 \cdot 2 \cdot 1 = 24\]Let’s try a sample problem so you can see how this works.
A salad bar has nine kinds of ingredients. Customers may mix four of the ingredients into their salad. How many possible ingredient combinations can customers come up with?
Solution
You’re looking for combinations in which \(9\) is the total number of ingredients and \(4\) is the number of ingredients that can be chosen. So, insert those numbers into the formula:
\[C = \frac{9!}{4!(9-4)!} = \frac {9!}{4!5!}\]Now, here you could just do the factorials as presented, but the numbers will get really big and it will take extra time. Luckily, some of the factorials cancel out:
\[\frac {9!}{4!5!} = \frac {9 \cdot 8 \cdot 7 \cdot 6}{4 \cdot 3 \cdot 2 \cdot 1} = \frac {3\text{,}024}{24} = 126\]Permutation is a similar concept as combinations, except with permutations the order of the things matters. Think of a four-digit passcode on a cell phone. Each of the digits in the code will be from \(0\) to \(9\), but it’s not enough to just know all four digits, you need to know their order. Permutations tell you how many possible passcodes you could have.
There are two formulas for computing permutations depending on whether repetition is allowed or not. When repetition is allowed, this is the formula:
\[P = n^r\]where \(n\) is the number of things to choose from and \(r\) is the number of choices.
When repetition is not allowed, this is the formula:
\[P = \frac{n!}{(n-r)!}\]Let’s try an example problem.
How many three-letter combinations can be formed from the letters a, b, c, d, e if a letter can’t be repeated in a single combination?
Solution
Here the order matters: abc is different from bac. Also, we can’t repeat letters, so aab isn’t permissible. As such, we’ll use the second permutation formula:
\[P = \frac{n!}{(n-r)!} = \frac{5!}{5-3!} = \frac{5 \cdot 4 \cdot 3 \cdot 2 \cdot 1}{2 \cdot 1} = 5 \cdot 4 \cdot 3 = 60\]BONUS: Problem Solving Tips
Much of the TABE allows you to use a calculator, so it is not testing your computing skill. Rather, you will be assessed on how well you can take given numbers and design a means for solving the problem, including determining which operation(s) to use and in what order to use them.
Here we are going to review some essential concepts and procedures with which you will need to be fluent. We have just given the basic information about them for you to use with our practice questions.
Skills to Practice
Many students cringe at the thought of word problems, and the math section of the TABE is made up of these. Not only do you have to do a given operation with numbers, but you first have to figure out which operation, or operations, to do. But there are clues in every word problem to help you out.
Choosing the Right Operations
Be familiar with key words and other clues, because these will make it easier for you to translate words and phrases into numbers and mathematical operations.
There are certain words and phrases that will let you know that you’re dealing with an equality or an inequality:
-
The words is and equal to suggest an equality of two statements and are represented by the equal sign (\(=\)).
-
Greater than and less than suggest an inequality, and are represented by the greater than sign (\(>\)) and less than sign (\(<\)), respectively.
-
Then there are the phrases greater than or equal to and less than or equal to, which are written mathematically as \(\ge\) and \(\le\), respectively.
Furthermore, watch out for these words and the operations they indicate:
-
Plus, combine, total, sum, together, more, and increase are used to indicate addition.
-
Minus, less, difference, left, take away, decreased, and fewer are used to indicate subtraction.
-
Times, product, twice, thrice, and of are used to indicate multiplication.
-
Per, quotient, and ratio are used to indicate division.
Identifying Unimportant Information
Irrelevant information may be deliberately inserted in math questions to make simple questions seem complicated or to mislead test-takers. Familiarity and constant practice in solving word problems will help you find what’s useful information in a math question.
There are also some useful tricks to make it easier to spot the unnecessary information. It helps to know formulas. A given number could be irrelevant if it does not fit in the formula. You should also sketch and label as you read the question. With a visual guide, an irrelevant piece of information will be easier to spot.
Manipulating Formulas
Some questions may require you to think about the formulas you’ve learned in a different way. For instance, they might provide the areas or volumes of a shape and ask for the other factors instead. The process is still the same: plug in the known values, represent the unknown values with letters, isolate the unknown value to one side of the equation, and solve for the unknown.
For example, a TABE math question may ask for the height of a cylindrical tank if it has a capacity of \(30\) cubic meters of water and has a radius of \(1.5\) meters.
Here, you are given the capacity (volume) of the tank and need the height. You can still use the volume formula you used before, you’ll just need to move things around:
\[V = \pi r^2 h\] \[30 = \pi (1.5)^2 h\]Isolate the unknown value and perform the required operations to get your answer:
\[h = 30 \div \pi(1.5)^2 = 4.2 \text{ m}\]All Study Guides for the TABE Test are now available as downloadable PDFs