Summary Information from Vectors by Groups
One of the most important and useful vector functions to master is tapply. The ‘t’ stands for ‘table’ and the idea is to apply a function to produce a table from the values in the vector, based on one or more grouping variables (often the grouping is by factor levels). This sounds much more complicated than it really is:
data<-read.table("c:\\temp\\daphnia.txt",header=T) attach(data) names(data) [1] "Growth.rate" "Water" "Detergent" "Daphnia"
The response variable is Growth.rate and the other three variables are factors (the analysis is on p. 479). Suppose we want the mean growth rate for each detergent:
tapply(Growth.rate,Detergent,mean)
BrandA BrandB BrandC BrandD
3.88 4.01 3.95 3.56
This produces a table with four entries, one for each level of the factor called Detergent. To produce a two-dimensional table we put the two grouping variables in a list. Here we calculate the median growth rate for water type and daphnia clone:
tapply(Growth.rate,list(Water,Daphnia),median)
Clone1 Clone2 Clone3
Tyne 2.87 3.91 4.62
Wear 2.59 5.53 4.30
The first variable in the list creates the rows of the table and the second the columns. More detail on the tapply function is given in Chapter 6 (p. 183).
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access