My three indoor cats are all seniors now, so I’m even more concerned that they have a healthy diet. I feed them both wet and dry food. They share most of two cans of wet food per day and have dry food, kibble, available all the time. Kibble is like snack time. Fortunately, they are more disciplined than I am; they are all of normal weight despite having access to all that food.
With their advanced age, I was concerned that the kibble be easy on their digestive systems. I had been feeding them Iams for the past six years since I rescued them and they seemed content with it. Nevertheless, I went to Chewy to see if they had any alternatives for senior cats with sensitive stomachs. To say the least, I was quite surprised that Chewy had so many brands catering to what I thought was a limited feline demographic. So much for an exhaustive experiment I envisioned, this was going to be a pilot test limited to just a few brands.
If this were a critical scientific experiment that would face peer review and replication, there would be a lot to consider in planning. But this is just a small, personal experiment between just me and my cats. We’ll all be satisfied with the results whatever they may be. So, with that said, here’s the outline of the experiment.
Population
The population for the experiment is small; it’s just my three indoor cats – Critter. Poofy, and Magic. Critter (AKA Ritter Critter) was rescued by my daughter from outside the Ritter building at Temple University when she was a kitten in 2006. Poofy (AKA Poofygraynod) was rescued from my back yard in 2013 when (the vet thought) she was about 6. Magic (AKA Black Magic) was also rescued (with Poofy) from my back yard in 2013 when (the vet thought) she was about 2. Critter and Magic were born in the wild; Poofy might have been a stray. At the time of the experiment, Critter and Poofy were about 14 and Magic was about 10. Poofy weighs about 13 pounds, Critter weighs about 11 pounds, and Magic weighs about 8 pounds, Poofy eats mostly kibble but also some wet food. Critter eats mostly wet food (ocean whitefish is her favorite) but also some kibble. Magic eats both kibble and wet food equally, in good feral fashion.

Because the population consists of only three individuals, and it’s their composite response that is of interest, the experiment actually involves measuring the cats’ preferences repeatedly over the course of the experiment. This is a census (rather than a survey) of the population. The sampling design is systematic, one set of measurements of kibble eaten by brand for the duration of the experiment. This is called a repeated measures design.
Phenomenon
The phenomenon being evaluated is preference for selected brands of kibble. Each cat may have a different preference, even changing day-by-day, but only the composite preference is important because I purchase the kibble in the aggregate. Preference is measured by the amount of each brand of kibble consumed in a 24-hour period.
Research Questions
I had four research questions I wanted to answer.
- What kinds of kibble do my indoor cats like to eat?
- I hypothesized that they might prefer Iams since that is the brand they had been eating for the prior seven years.
- I hypothesized that they might like seafood best because this preference is often depicted in cat-related cliche.
- Protein, Fat, Fiber. I hypothesized that they might prefer the highest protein content.
- I hypothesized that they might like kibble that was smaller and more rounded so that it was easier to swallow.
- How much do they eat in a day? I hypothesized that they would eat less that three cups of food per day based on prior feeding patterns. I provided a total of about twice that amount during the experiment, one cup of kibble for each of the six brands each day.
- Do they prefer variety or will they eat the same kibble consistently? I hypothesized that they would eat a variety of the brands because I like variety in MY diet. My vet disagreed. He thought they would eat what they were most familiar with.
- Will a different kibble reduce their barfing? I hypothesized that there would be no difference because they were already eating Iams kibble for cats with sensitive stomachs.
Kibble Brands
I was surprised by how many different brands there are of dry food for senior cats with sensitive stomachs available from just one vendor (Chewy). I decided to test just six because of cost and test logistics. They were:
Iams – because that’s the brand I had been feeding the cats for the last seven years; it was my “control group” brand. It is the least expensive of the brands and consists primarily of chicken and turkey, corn, rice, and oats.

Purina – because I wanted a well-known national brand available in supermarkets. I selected two Purina products, Focus and True Nature, to see if there was a difference between kibble formulations from the same company. They have the largest kibble, high protein, high calories, and tend to be more expensive. Focus is chicken and turkey flavored and contains rice, oats, and barley. True Nature is salmon and chicken flavored, and is grain free.


Hills Science Diet – because I wanted a well-known, highly-rated, vet-recommended brand sold mostly in pet stores. It is high density and high calorie but lower protein than the other brands. It is chicken flavored and contains corn, soy, and oats.

Halo – Perhaps the first “holistic” cat food, having been introduced in the 1980s, it purports to be ultra-digestible because of its use of fresh meats, vitamins, probiotics, and other healthy ingredients. It is seafood flavored and contains oats, soy, and barley.

I and Love and You – A newer formulation of holistic ingredient, it is grain-free, includes prebiotics and probiotics for healthy digestion, and has the longest ingredient list by far. It contains seafood and chicken/turkey.

Experimental Procedure
The experimental set up consisted of six paper bowls, one for each brand of kibble. The positions of the bowls were randomized so that the cats wouldn’t associate a certain kibble brand with a position. Every 24 hours, at 8 PM, the bowls were filled with one cup of kibble and weighed. (They still got their can of wet food at 5 AM when they wake me up.) The cats were then allowed to eat the kibble as they wanted for 24 hours. It was clear from the beginning of the experiment that the cats did prefer certain brands, though they would try others.

At the end of 24 hours, the bowls were reweighed. Remaining kibble was transferred to a bucket to be fed to my three outdoor feral cats, who will eat anything. The bowls were then filled and weighed again, and placed in their new randomized position.
Data Recorded each day included:
- Day of experiment and time
- Kibble brand
- Bowl position
- Weight of kibble not eaten
- Weight of kibble provided for the next day
An informal inspection of the house was also conducted to identify any barfs that may have occurred.

Planned Analysis
The dependent variable for the analysis was the brand of kibble. The independent variable for the analysis was the weight of kibble eaten, calculated as:
Weight of kibble eaten = Weight of kibble provided – Weight of kibble left over
The position of the bowls was a blocking factor used to control extraneous variance. The day-of-the-experiment was a repeated-measures factor. This design is a two-way repeated-measures Analysis of Variance (ANOVA).
Prior to conducting any statistical testing, an exploratory analysis was planned involving calculating descriptive statistics and constructing graphs.
Depending on the results of the exploratory analysis, global ANOVA tests were planned for detecting differences between the brands, with the effects of bowl position and day of the experiment held constant. A priori tests were also planned to detect any differences between individual brands and the control brand, Iams.

Results
Though not what I expected, it was obvious after a couple of days that the cats had a clear preference for Hills Science Diet. Consequently, I ended the experiment after two weeks.

Descriptive Summary
The following table summarizes the amount of each brand that the cats ate over the two-week experiment.

Statistical Testing
While the design of the experiment is technically a two-way repeated-measures Analysis of Variance (ANOVA), the large differences in brands and the lack of differences in bowl position and day of the experiment made calculating the model unnecessary. This solved the problem of my not having appropriate software to conduct that part of the analysis. Instead, the following sections describe two-way ANOVA results for the brand versus bowl position and the brand versus day of the experiment models. Statistical comparisons of the amounts of each brand eaten are also summarized, with emphasis on differences between each brand and the “control” brand, Iams.



The global ANOVA test for the brands was significant when the effects of the day of the experiment was controlled for. The day of the experiment had no impact on the amount of kibble eaten. No surprise there.




The global ANOVA test for the brands was significant when the effects of bowl position was controlled for. Bowl position had no impact on the amount of kibble eaten. Again, that’s not a big surprise although you can never be sure of a hypothesis when you’re working with cats.


The following table summarizes the statistical tests between the brands. The important tests are the comparisons between each brand and the control brand, Iams, highlighted in yellow. The only significant tests were the comparisons between Hills Science Diet versus Iams and Purina True Nature versus Iams. This means that my cats like Hills and True Nature a lot more than what I’ve been feeding them for the last few years. Time to switch brands. I could have done worse had I fed them one of the other brands instead of Iams, but not significantly so.


Findings
First, things don’t always go the way you think they will. This is true in any experiment … and life in general. There was really no need to conduct the sophisticated ANOVA that I had planned, so I didn’t bother. Oh well, next time.
Second, my three cats eat about 135 grams (4.8 oz) of kibble in a day. Now I can use the automatic reorder feature on Chewy and save a few dollars.
Third, You know how you tend to eat a lot more after you come home with new groceries? Cats do it too. They ate a lot more on the first day of the experiment when they had five new brands of kibble to taste.
Fourth, I randomized bowl position as a way of controlling extraneous variation for the ANOVA. It seems that the middle positions had more kibble eaten per bowl than the outer positions. I have no explanation for this pattern and the cats aren’t talking.
Fifth, I can’t say that it reduced barfing because I had no baseline. My daughter, who doesn’t like cats, led me to believe that they barf about every twenty minutes. Still, there were only three barfs during the two-week experiment, which I considered not to be so bad.
Sixth, my cats clearly prefer Hill’s Science Diet as their kibble of choice. They don’t seem to want a variety of brands. I don’t know why my cats preferred Hills. It doesn’t appear to be the flavor, texture, or protein content since other brands had different combinations of these factors. If you look at reviews of other brands of kibble, you’ll find people who swear that their cat(s) likes the-brand-that-they-buy best. They’re probably right. Every cat or population of cats may have different tastes. I, myself, like pineapple on my pizza.
Finally, my experiment wasn’t large or sophisticated enough to isolate and analyze hypotheses about ingredients. If I could figure out why cats prefer one brand of kibble over another, though, I could probably get a job with Purina.

Further Research
It’s always a good practice to describe additional research that could be done to make the world a better place. Who knows if somebody with money might see it and fund your further research? In this case, further research might involve testing different brands, especially if the brands could be selected to explore a variety of flavors, kibble shapes, sizes, densities, and types and concentrations of protein. Finally, I would recommend using many, many more cats if you can. My daughter won’t let me have any more.
So, if you find yourself with some time on your hands, consider conducting your own experiment on your cats. You might be surprised at what you learn. It’ll be fun.

Read more about using statistics at the Stats with Cats blog. Join other fans at the Stats with Cats Facebook group and the Stats with Cats Facebook page. Order Stats with Cats: The Domesticated Guide to Statistics, Models, Graphs, and Other Breeds of Data analysis at amazon.com, barnesandnoble.com, or other online booksellers.
What to Look for in Data – Part 1 discusses how to explore data snapshots, population characteristics, and changes. Part 2 looks at how to explore patterns, trends, and anomalies. There are many different types of patterns, trends, and anomalies, but graphs are always the best first place to look.
There are four patterns to look for:
Simple trends can be easier to identify because they are more familiar to most data analysts. Again, graphs are the best place to look for trends.
Linear trends are easy to see; the data form a line. Curvilinear trends can be more difficult to recognize because they don’t follow a set path. With some experience and intuition, however, they can be identified. Nonlinear trends look similar to curvilinear trends but they require more complicated nonlinear models to analyze. Curvilinear trends can be analyzed with linear models with the use of
Temporal
Spatial Trends present a different twist. Time is one-dimensional; at least as we now know it. Distance can be one-, two-, or three-dimensional. Distance can be in a straight line (“as the crow flies”) or along a path (such as driving distance). Defining the
Categorical Trends are no more difficult to identify than any trend except you have to break out categories to do it, which can be a lot of work. One thing you might see when analyzing categories is 
Anomalies



Some activities are instinctive. A baby doesn’t need to be taught how to suckle. Most people can use an escalator, operate an elevator, and open a door instinctively. The same isn’t true of playing a guitar, driving a car, or analyzing data.
Scrubbing your data will make you familiar with what you have. That’s why it’s a good idea to know your objective first. There are many things you can do to scrub your data but the first thing is to put it into a matrix. Statistical analyses all begin with matrices. The form of the matrix isn’t always the same, but most commonly, the matrix has columns that represent variables (e.g., metrics, measurements) and rows that represent observations (e.g., individuals, students, patients, sample units, or dates). Data on the variables for each observation go into the cells. Usually this is usually done with spreadsheet software.
What To Look For
Snapshots aren’t difficult. You just decide where you want a snapshot and record all the variable values at that point. There are no descriptive statistics, graphs, or tests unless you decide to subdivide the data later. The only challenge is deciding whether taking a snapshot makes any sense for exploring the data.

For properties graphs (bar charts, area charts, line charts, candlestick charts, control charts, means plots, deviation plots, spread plots, matrix plots, maps, block diagrams, and rose diagrams), look for the unexpected. Are the central tendency and dispersion what you might expect? Where are big deviations?
Part 3 of Dare to Compare shows how one-population statistical tests are conducted. Part 4 extends these concepts to two-population tests.


This is a bit more complicated than the formula for a one-population test because you can have different standard deviations and different numbers of measurements in the two populations.
If the number of measurements taken of the two populations is the same, the test design is said to be balanced. If the variances of the measurements in the two populations are the same, the leftmost term in the denominator reduces to s2. So, the formula for a balanced two-population t-test with equal variances is:
Much more simple but not as useful as the more complicated formula. You might be able to control the number of samples from the populations but you can’t control the variances.



So there is no statistically significant difference between the Fall semester classes and the Spring semester classes.

Now for the days of the week:

Here is a summary of the three tests.
So what do you do if you have more than two populations or more than one phenomenon or some other weird combinations of data? You use an Analysis of Variance (ANOVA).
Multi-Way ANOVAs
Random Effects ANOVAs assume that the levels of a main effect are sampled from a population of possible levels so that the results can be extended to other possible levels. The Instructors main effect in the example could be a random effect if other instructors were considered part of the population that included Dr. Statisticus and Prof. Modearity. If only Dr. Statisticus and Prof. Modearity were levels of the effect, it would be called a fixed effect. If a design included both fixed and random effects, it is called a mixed effects design.
Parts 1 and 2 of Dare to Compare summarized fundamental topics about simple statistical comparisons. Part 3 shows how those concepts play a role in conducting statistical tests. The importance of these concept are highlighted in the following table.





In this comparison, the table t-value you would use is for a one-tailed (directional) test at 90% confidence for 10 samples, t(1-tailed, α = 0.1, 9 degrees of freedom) = 1.383. For comparison, the value of t(2-tailed, 0.9 confidence, 9 degrees of freedom), which was used in the first example, is equal to 1.833, as is t(1-tailed, 0.95 confidence, 9 degrees of freedom). The reason is that you only have to look in half of the t-distribution area in a one-tailed test compared to a two-tailed test. That means that if you use a directional test you can have a smaller false positive rate.


The American Statistical Association has identified 
Part 1 of Dare to Compare summarized several fundamental topics about statistical comparisons.
Statistical tests that don’t rely on the distributions of the phenomenon in the populations are called nonparametric tests. Nonparametric tests often involve converting the data to ranks and analyzing the ranks using the median and the range.
When you conduct a statistical test, the result does not mean you prove your hypothesis. Rather, you can only reject or fail to reject the null hypothesis. If you reject the null hypothesis, you adopt the alternative hypothesis. This would mean that it is more likely that the null hypothesis is not true in the populations. If you fail to reject the null hypothesis, it is more likely that the null hypothesis is true in the populations.
After you conduct the test, there are two pieces of information you need to determine – the sensitivity of the test to detect differences, called the effect size, and the power of the test. The power of the test will depend on the sample size, the confidence, and the effect size. The effect size also provides insight into whether the test results are meaningful. Meaningfulness is important because a test may be able to detect a difference far smaller than what might of interest, such as a difference in mean student heights less than a millimeter. Perhaps surprisingly, the most common reason for being able to detect differences that are too small to be meaningful is having too large a sample size. 
z-Tests and t-Tests
χ2 Tests


In school, you probably had to line up by height now and then. That wasn’t too difficult. There weren’t too many individuals being lined up and they were all in the same place at the same time. An individual’s place in line was decided by comparing his or her height to the heights of other individuals. The comparisons were visual; no measurements were made. Everyone made the same decisions about the height comparisons. You didn’t need statistics to solve the problem. So why might you ever need statistics to compare heights?
Fortunately, you don’t have to measure every individual in the population so long as you measure a representative sample of the individuals in the populations. You can improve your chances of getting a representative sample by using the three Rs of variance control —
A bell curve is usually assumed to represent a Normal distribution. The average and the variance of the values are called parameters of the distribution because they are in the mathematical formula that defines the form of the distribution.
Read more about using statistics at the
Whether you know it or not, you deal with models every day. Your weather forecast comes from a meteorological model, usually several. Mannequins are used to display how fashions may look on you. Blueprints are drawn models of objects or structures to be built. Maps are models of the earth’s terrain. Examples are everywhere.

Models can also be expressed in words and pictures. These are used in virtually all fields to convey mental images of some mechanism, process, or other phenomenon that was or will be created. Blueprints, flow diagrams, geologic fence diagrams, anatomical diagrams are all conceptual models. So are the textual descriptions that go with them. In fact, you should always start with a simple text model before you embark on building a complex physical or mathematical model.
Theoretical Models
Probability Models
In statistical comparison models, the dependent variable is a grouping-scale variable (one measured on a nominal
Clustering models do not have a nominal-scale dependent variable, but most classification models do.
Factor Analysis
Some models are created to predict new values of a dependent variable or forecast future values of a time-dependent variable. To be useful, a prediction model must use prediction variables that cost less to generate than the prediction is worth. So the predictor variables and their scales must be relatively inexpensive and easy to create or obtain. In prediction models, accuracy tends to come easy while precision is elusive. Prediction models usually keep only the variables that work best in making a prediction, and they may not necessarily make a lot of conceptual sense.
Step 1 – Start at top of the Catalog of Models figure. Decide whether you want to create a physical, mathematical, or conceptual model. Whichever you choose, start by creating a brief conceptual model so you have a mental picture of what your ultimate goal is and can plan for how to get there.
Say you wanted to describe someone you see on the street. You might characterize their sex, age, height, weight, build, complexion, face shape, hair, mouth and lips, eyes, nose, tattoos, scars, moles, and birthmarks. Then there’s clothing, behavior, and if you’re close enough, speech, odors, and personality. Your description might be different if you’re talking to a friend or a stranger, of the same or different sex and age. Those are a lot of characteristics and they’re sometimes hard to assess. Individual characteristics aren’t always relevant and can change over time. And yet, without even thinking about it, we describe people we see every day using these characteristics. We do it mentally to remember someone or overtly to describe a person to someone else. It becomes second nature because we do it all the time.


Shape refers to the frequency of the values in a dataset at selected levels of the scale, most often depicted as a graph. For ordinal scales, the graph is usually a histogram. For continuous scales, the graph is usually a probability plot, although sometimes histograms are used. Shapes of continuous scale data can be compared to mathematical models (equations) of frequency distributions. It’s like comparing a person to some well-known celebrity; they’re not identical but are similar enough to provide a good comparison. There are dozens of such distribution models, but the most commonly used is the 




