CENSA User Guide

How to select and arrange Census data

Published

Saturday Jun 28, 2025 at 10:32 PM EDT

1 Introduction

1.1 Purpose

This guide provides instructions for using the CENsus Select and Arrange (CENSA) web application. CENSA allows epidemiologists, analysts, and other IDOH users to select standardized population data from the US Census Bureau and modify the format to best suit their analysis or reporting needs. Once the user finalizes their selections and formatting, they can download an Excel file that can be shared with other collaborators.

1.2 Usage overview

To access the app, load https://on.in.gov/censa in your web browser while connected to the State of Indiana’s public network.

The CENSA workflow consists of the following steps:

  1. Filter the data by vintage type, geography level, and population year to get a smaller set of data
  2. Select demographics, including age group, sex, ethnicity, and race
  3. Arrange the output
  4. Download an Excel file
  5. Repeat (if necessary), for other geographies or combinations of demographics

Each step is explained in the sections below, with annotated screen shots from the app. Key features in the screen shot are marked by numbers; refer to the corresponding number in the text description or below the image for an explanation of the feature.

Click images to expand

Click the images in this document to see an enlarged version. This works on the HTML version but not the PDF.

1.3 Getting help

There are several resources available if you need help:

2 Filter

The full Census data is quite large, so as a first step it is important to apply some filters to reduce the data. Refer to Figure 1.

Figure 1: Filter selections

The first selection has been made for you—the ‘Best’ vintage (1) will always be displayed as this is the best option for most users.

Using other vintages

Refer to the Census FAQ document for more information on available vintages. To use a vintage other than the one recommended by ODA, you’ll need to use SQL. You can switch to the Query tab in CENSA, write your own SQL, and comment out the line that sets vintage_best = 1. An alternative is to get access to the ACE Synapse database and query the data there.

Tool tips

Hover over the icons for more guidance and information pertaining to that specific section. Figure 2 below shows an example of an expanded tooltip.

Figure 2: Hover over the icon to see a tooltip

Next, select a geography level (2). There are nine geographies to choose from:

  1. nation
  2. region
  3. division
  4. state
  5. district_care_coordination
  6. district_public_health
  7. district_sti
  8. district_zero_is_possible
  9. county
More info about districts

Refer to the the Census FAQ document for more information about the IDOH districts listed above.

Finally, select one or more population years to include (3), and then click the ‘Apply filters’ button (4). CENSA will load the filtered data and display the results in the table in the main section of the page as shown in Figure 3 below.

Select fewer years for better performance

CENSA allows you to select multiple population years, but for best performance (particularly for more granular geography levels like counties) select fewer population years (3) at a time to reduce the amount of data.

Remember to click the ‘Apply filters’ button (4)

Let’s say you have selected and applied the filters described in this section (Section 2), and then you have made some demographic selections in the next section (Section 3). You realize you want to add another population year, or maybe you want to change the geography level. Once you have made those changes to the filters in this section, you must click the ‘Apply filters’ button (4) again. This may update or reset the values available in the demographic selections, because they depend on the selected filters.

3 Select demographics

You may continue to refine the data needed for your use case by selecting the combinations of demographic values from the filters in the left panel. Available demographic values can change based on the geography level selected. The national and state geography levels have the most age groups to choose from, while the district and county will have a more limited selection of standard age groups (or age groups that can be aggregated from the standard age groups). More information about the availability by geography level and demographics can be found in the Census FAQ document.

Refer to Figure 3 below. You may choose demographic selections in any order. For example, you can go in order from top to bottom, selecting values from the age group dropdown (1) followed by the sex (2), ethnicity (3), and race (4) dropdowns. Or you can start with race, then select age groups, etc.

Figure 3: Select demographic values
Note

The values available in one demographic filter may depend on the values previously selected in other demographic filters.

3.1 Race category best practices

The US Census Bureau uses two methods to categorize individuals by race. For more information regarding these categories, hover over the tooltip beside the Race selection in the app (as shown below in Figure 4) or refer to the Census FAQ.

Figure 4: Race tooltip

3.2 Combining ethnicity and race

Some users may wish to combine ethnicity and race, with categories for Non-Hispanic American Indian or Alaska Native, Non-Hispanic Asian, etc., and Hispanic (regardless of race). If this isn’t relevant, skip to Section 4.

Warning

Ethnicity and race are separate concepts, and it is recommended to keep them separate.

To achieve this combination, begin by making the making the following selections in CENSA (as shown in Figure 5):

  1. Select ‘Hispanic’ and ‘Not Hispanic’ from the Ethnicity dropdown
  2. Select ‘All Races’, the five ‘alone’ race values (refer to Section 3.1), and ‘Two or More Races’ from the Race dropdown (this is the default set of race selections when you first load CENSA)
  3. Make sure ‘ethnicity’ and ‘race’ have the box next to them checked in the Arrange section, to include those values as rows
Figure 5: Demographic selections for combining ethnicity and race

After making those selections, download the Excel file, open it, and go to the population sheet. In Excel, make the selections as shown in Figure 6:

  1. Use the mouse and the shift key to select the rows where ethnicity is ‘Hispanic’, except where race is ‘All Races’
  2. Also select the row(s) where ethnicity is ‘Not Hispanic’ and race is ‘All Races’ (using the ctrl key on Windows or command on macOS)
  3. Right-click the selection, and click ‘Hide’
Figure 6: Hide some combinations of ethnicity and race in Excel

The result should look like Figure 7.

Figure 7: Ethnicity and race combined in Excel

4 Arrange

4.1 Ordering rows and columns

Check or uncheck the boxes next to the variables in the Arrange section (Figure 8) to dynamically format the population data to suit your specific needs.

Figure 8: Table that identifies how to arrange variables and their values

Every variable that needs to be arranged based on the previous filters and demographic selections is automatically included in this table. If the box next to a variable is checked, the variable’s name will appear as a column in the output table and its unique values that have been selected will appear as separate rows; this is referred to as a row variable. Uncheck the box, and the variable becomes a column variable, where each unique value will appear as a separate column.

Example

If the box next to ethnicity is checked, the output table to the right will have an ethnicity column with a row for ‘Hispanic’ and another for ‘Not Hispanic’. If it is unchecked there will be only one row for that data, with a column called hispanic and another called not_hispanic.

By checking or unchecking the variables, users can create different layouts:

The images below show examples of the various arrangements. Row variables are represented by a teal arrow and column variables by an indigo arrow.

4.1.1 Long format (all variables are row variables)

Check the boxes next to all the arrangeable variables to achieve an output table with a long format, as shown in Figure 9.

Figure 9: Long Format
Note

If only a single value is selected for a demographic or other filter, that variable will be automatically treated as a row variable and will not appear in the table to dynamically arrange.

For example, in Figure 9 the user has selected to only show the population for Hispanics. As a result, the ethnicity variable does not appear in the Arrange table but is instead shown as a row variable—the population table has a column name of ethnicity with row values of ‘Hispanic’.

4.1.2 Wide format (all variables are column variables)

Uncheck all the arrangeable variables to achieve an output table with a wide format, as shown in Figure 10. When there are multiple column variables, their values get combined, converted to lower case, and separated by a pipe. For example, the column named age_0-4|american_indian_or_alaska_native_alone|year_2022 represents the population for those age 0-4 of race ‘American Indian or Alaska Native Alone’ for the year 2022.

Figure 10: Wide format

4.1.3 Hybrid format (some are row variables, one is a column variable)

Figure 11 shows a hybrid example, with multiple row variables (race and population_year) and one column variable (age_group).

Figure 11: Hybrid format

4.1.4 Hybrid format (one is a row variable, two are column variables)

Figure 12 shows another hybrid example, with one row variable (population_year) and multiple column variables (age_group and race).

Figure 12: Another hybrid format

4.2 Sorting rows

Click on a column header in the output table to sort the rows of that column.

Sorting tip

The output table in the CENSA app is meant for simple visual checks, and only allows sorting by one column at a time. If it is important to sort by multiple columns or use custom sort orders, please export the data and apply the sorting in the downloaded Excel file as shown in Figure 16.

5 Download

After filtering the data, selecting demographics, and arranging the columns, the final step in the CENSA flow is to download the data as an Excel file. Click the ‘Prepare arranged data’ button (1), enter your name (2) and email address (3), and click the ‘Submit and download’ button (4) as shown in Figure 13. The file should download, and then you can close the popup by clicking on the x in the top right corner (5).

Figure 13: Prepare and download an Excel file with population data
Note

Depending on the filters, selections, and arrangements, the output table displayed in CENSA may be truncated. However, the Excel file will still include all the rows.

5.1 Excel sheets

The downloaded Excel file contains five sheets, which are briefly described in the following sections.

5.1.1 notes

The notes sheet describes the data, and provides caveats for responsible use.

Figure 14: The notes sheet

5.1.2 selections

The selections sheet captures the settings selected in the app that generated the output Excel file, and the dates the data and app were updated.

Figure 15: The selections sheet

5.1.3 population

The population sheet contains population data for use in analysis, reports, and dashboards. Use the features of Excel to apply more advanced sorting to the rows of the population table, if desired.

Figure 16: The population sheet
Note

For each population value in the sheet, there is a corresponding value in the same cell in the source_row (Section 5.1.4) and vintage (Section 5.1.5) sheets. For example, a population in cell B6 has corresponding values in B6 on the source_row and vintage sheets that refer to it.

5.1.4 source_row

The source_row sheet identifies the source file(s) from the US Census FTP site and the exact row number(s) used to generate the population value, for traceability. If it was the result of an aggregation or calculation, the source value will have multiple source/row identifiers. If it has not been transformed in any way, it will have a single source/row identifier.

Figure 17: The source_row sheet

5.1.5 vintage

The vintage sheet identifies the Census vintage, and is helpful for traceability.

Figure 18: The vintage sheet

6 Repeat (if necessary)

If you need data for multiple geography levels, perform the entire workflow for the first geography and download the resulting file. Then, update the geography level, make any additional changes in the filter section, click the ‘Apply filters’ button, tweak any other demographic selections if necessary, and perform the download process again.

Note

To reset all the filters and demographic selections, refresh the page in your browser.

7 Population, Source, Vintage, Query, and Documentation tabs

Figure 19 shows the five navigation tabs near the top of the CENSA application.

Figure 19: Population, Source, Vintage, Query, and Documentation tabs

The Population tab is the one most users will use most of the time. It shows the population data and is updated based on the filters, demographic selections, and data arrangements selectors in the left panel.

The Source and Vintage tabs have the same format as the Population tab, but show the identifiers of the US Census data source files and the vintages selected, for traceability purposes.

The Query tab is for power users who prefer to write their own SQL queries using the standardized population data that powers CENSA.

The Documentation tab has previews and links to this CENSA User Guide and to the Census FAQs.

Figure 20: Documentation tab

8 Troubleshooting

Refer to this section for troubleshooting tips if you encounter issues with CENSA.

8.1 CENSA container app is stopped after business hours

To conserve costs, the Azure container application that powers CENSA is shut down on the weekends and from 10pm until 4am during the week. If you try to use CENSA during those times you will see an error page like Figure 21. Please wait until business hours and try again.

Figure 21: CENSA will show this error page when it is stopped after business hours

8.2 CENSA is slow to load

If no one has used CENSA recently, the servers hibernate during business hours to save costs. If you are loading the application in your browser for the first time since it hibernated, you will need to wait slightly longer (roughly 20 seconds) for CENSA to be ready.

8.3 Unable to load CENSA

CENSA was developed for the Indiana Department of Health, and the website is blocked unless you access it from the State of Indiana public subnet. If you get an error message RBAC: access denied when loading https://on.in.gov/censa, connect to the public state network at an IDOH office, via the IOT VPN, or from within the IDOH ACE environment (if you have access to it).

8.4 CENSA is stuck or gives an error

If the app becomes stuck when loading or the Query tab displays an error, clear your browser’s cache and refresh the page.

8.5 Too much data to download

Microsoft Excel can only handle a maximum of 1,048,576 rows and 16,384 columns. You might exceed these limits if you select a small geography level (such as county) with several years of data, depending on how you have arranged the variables. If you get a message about too many rows or columns, select less data or rearrange the columns before attempting to download it again. For example, select fewer years or demographic values. Convert some variables to row variables if you have too many columns, or to column variables if you have too many rows.

8.6 Download error: “file wasn’t available on site”

If your browser’s download icon shows an error like Figure 22 after you hit ‘Download arranged data’, click the button to download the data again. It should work the second time.

Figure 22: Error indicating the file should be downloaded again

8.7 Problem with content in the Excel file

When you try to open the downloaded Excel file, you may see an alert like the one shown in Figure 23 that mentions a problem with the content. Click ‘Yes’ and the file should open correctly.

Figure 23: Click ‘Yes’ to recover the content in the downloaded Excel file

8.8 I need data that is not available in CENSA

CENSA includes population data for many geographies, population years, and demographic combinations, but it is not exhaustive. Some data is just not available from the US Census. Other data is available, but in a different format that ODA has deemed a lower priority to include. For example, CENSA does not contain data by ZIP code. The Census FAQ document contains additional details.

In the future, ODA may incorporate additional Census data. If you have a suggestion, please send an email to Jamie Black (jamblack@health.in.gov).

9 Feedback

The Office of Data and Analytics hopes CENSA makes it easier and more efficient for IDOH team members to incorporate population data in analyses and reports so that they can repurpose that time to positively impact the lives of Hoosiers. Was CENSA helpful? Or was your experience with CENSA less than ideal? Whatever the case, we want to hear your feedback! Let us know what does or doesn’t work well, about bugs you’ve encountered, and ideas for improvements or features you would like to see in the future. Click the menu icon (three dots) in the upper-right corner of the app as shown below to submit your feedback.

Figure 24: How to get more info and submit feedback