Title
desctable --
Tables of Descriptive Statistics - with Unique Treatment of
Continuous, Binary, and Nominal Variables
General syntax
desctable varlist [if] [in] , filename( ) [options]
Overview
desctable produces a descriptive (summary) statistics table. desctable
treats continuous, binary, and nominal variables differently -- providing
statistics and labeled output most appropriate for the measurement level
of the variable.
Factor syntax is required: binary and nominal variables must be entered
into the variable list with the i. prefix. Variables without any prefix
are assumed to be continuous.
desctable outputs the descriptive statistics table to Excel; it can easily
be copied to Word without losing the formatting of the table.
desctable labels the rows of the table with the corresponding variable
label. Categories of nominal variables are labeled with the value labels
of the corresponding variable. Thus, the effectiveness of the default
table depends on the effectivness of the labels in the data (see label for
details on adding variable and value labels to your data).
Options
+-----------------+
----+ Required Option +---------------------------------------------------
filename(string)
is required. This names the Excel sheet that is saved with the
descriptive statistics table. If the filename includes spaces,
it must be enclosed in quotations; e.g. filename("file name
has spaces")
+---------------------------------------------+
----+ Optional statistics to include in the table +-----------------------
stats(string)
specifies which statistics should be included in the table.
Each statistic is reported in a separate column of the table.
The default is stat(mean sd) which reports the mean
[proportion for binary and nominal variables] and the standard
deviation of continuous variables. Statistics can be
requested in any order desired; columns of the table will be
ordered based on the list specified in stats( ). The following
statistics are available:
Stat Name Description
-----------------------------------------------------------
mean Mean of continuous variables; proportion of
binary/nominal variables.
sd Standard deviation (c).
freq Freqency for each category of a nominal
variable.
n Number of non-missing observations for a
variable.
variance Variance (c).
median Median (c).
min Minimum (c).
max Maximum (c).
range Range [max - min] (c).
iqr Interquartile range [75th percentile - 25th
percentile] (c).
cv Coefficient of variation [sd/mean] (c).
semean Standard error of mean (c).
skewness Skewness (c).
kurtosis Kurtosis (c).
p1 1st percentile (c).
p5 5th percentile (c).
p10 10th percentile (c).
p25 25th percentile (c).
p50 50th percentile [same as median] (c).
p75 75th percentile (c).
p90 90th percentile (c).
p95 95th percentile (c).
p99 99th percentile (c).
svymean *Survey weighted estimate of mean; proportion
of binary/nominal variables.
svysemean *Survey weighted estimate of standard error of
mean (c).
mimean *Multiple imputation estimate of mean;
proportion of binary/nominal variables.
misemean *Multiple imputation estimate of standard
error of mean (c).
misvymean *Survey weighted multiple imputation estimate
of mean; proportion of binary/nominal
variables.
misvysemean *Survey weighted multiple imputation estimate
of standard error of mean (c).
-----------------------------------------------------------
(c) indicates this statistic is only reported for continuous variables.
* svy statistics are only available if data has been svyset.
mi statistics are only available if data has been mi set.
misvy statistics are only available if data has been mi svyset.
+---------------+
----+ Other options +-----------------------------------------------------
decimals(#) changes the number of decimal places reported for the
statistics. The default is 2. Any integer between 0 - 5 is
allowed.
listwise performs listwise deletion across all specified variables.
That is, observations with missing data on any of the
variables included in the table are excluded from any of the
reported statistics. The default is casewise deletion which
includes observations that are non-missing for each specific
statistic.
notes(string)
adds footnotes to the bottom of the table. The notes are
automatically wrapped so they do not exceed the width of the
table. For long lists of notes, you will need to adjust the
height of the cell in Excel after the table is made for all of
the notes to be visible. Notes must be enclosed in quotation
marks, e.g. notes("This is a footnote"). For long lists of
notes you can list the notes on multiple lines in the code, as
long as each line is enclosed in quotations, e.g. notes("This
is the first footnote" "This is the second footnote"). The
appearance of the notes will be unaffected.
title(string)
adds a title to the table. The default is to title the table
"Table #: Descriptive Statistics (N = ##)" where the N size is
calculated automatically based on the total number of
observations that descriptive statistics were calculated for.
font(string) changes the font of all numbers, labels, notes, and titles in
the table. The default is "Times New Roman"
fontsize(#) changes the font size of all numbers, labels, and titles in
the table. The default is 11.
notesize(#) changes the font size of the footnotes on the table. The
default is 9.
txtindent(#) changes how far the statistics are indented from the right
edge of the column. All statistics are right-justified so that
the statistics align on the decimal point. The default is 1.
singleborder changes to a single horizontal line as the border on the top
and bottom of the table. The default is to use double
horizontal lines as the table top and bottom borders.
sheetname(string)
adds a name to the individual Excel sheet where the table is
saved. The default is to name the sheet "Descriptives Table"
varnames labels the rows of the table with the variable names rather
than the variable labels. The default is to use the variable's
label if it exists.
+---------------+
----+ Group options +-----------------------------------------------------
group(groupvar)
specifies that descriptive statistics should be calculated
separately for each group of the nominal grouping variable.
Statistics for each group are added to the right of the
tables. Labels for the new columns are automatically generated
from the value labels of the grouping variable.
Examples
sysuse nlsw88
desctable wage age i.race i.union i.collgrad tenure i.occupation hours,
filename("descriptivesEX1")
desctable wage age i.race i.union i.collgrad tenure i.occupation hours,
filename("descriptivesEX2") stats(mean freq sd min max iqr median)
desctable wage age i.race i.union i.collgrad tenure i.occupation hours,
filename("descriptivesEX3") font("Helvetica") fontsize(13)
desctable wage age i.race i.union i.collgrad tenure i.occupation hours,
filename("descriptivesEX4") decimals(3)
desctable wage age i.union i.collgrad tenure i.occupation hours,
filename("descriptivesEX5") group(race)
desctable wage age i.race i.union i.collgrad tenure i.occupation hours,
///
filename("descriptivesEX6") ///
notes("This is the first footnote." ///
"This is how to split long footnotes" ///
"onto multiple lines of the code.")
Comments
desctable makes the descriptive statistics table using Stata's putexcel
command. Because of some of the putexcel features used, Stata version 14.1
or newer is required to use desctable.
Available statistics for continuous variables are those that can be
calculated with tabstat (with the exception of q - which is not allowed
with desctable). See tabstat for more details.
Authorship
desctable is written by Trenton D Mize (Department of Sociology, Purdue
University) and Bianca Manago (Department of Sociology, Vanderbilt
University). Questions can be sent to tmize@purdue.edu
Back to top