Researchers and Advisors
What's the difference between these terms, and why does it matter?
Understanding these frequently mixed-up terms will help you develop a stronger protocol with better protections for study participants. Using correct terms in recruitment and consent materials contributes to building a relationship of trust between you and the potential participants. It also assists with creating an appropriate and effective data management plan.
Anonymity
Anonymity in research means no one—not even you, the researcher—knows who the participants are. A serious breach of confidentiality can occur when individuals provide responses while believing their identities can’t be known, yet certain information might directly or indirectly identify them. If such information were accidentally released without permission, that would violate participants’ trust and could put them at risk.
Often when researchers tell participants that their responses will be "anonymous," the researchers mean that they will use the data anonymously by removing names or other identifying information, using pseudonyms, or reporting data in the aggregate. If that's the case, that's what you need to tell individuals—not that the study or data is anonymous (unless even you don't know or can't ascertain someone's identity).
Even though the 1996 HIPAA (Health Insurance Portability and Accountability Act) refers specifically to medical or healthcare information, the WWU HRPP joins a majority of HRPPs and IRBs in using the 18 HIPAA identifiers and the HIPAA definition of PHI (Personal Health Information) to establish guidelines around identifiability. (Note that any student-related data is also subject to the Family Educational Rights and Privacy Act (FERPA).
HIPAA identifiers referenced in the federal regulations:
- Names;
- All geographical subdivisions smaller than a state, including street address, city, county, precinct, ZIP code, and their equivalent geo-codes, except for the initial three digits of a ZIP code, if according to the current publicly available data from the Bureau of the Census: (1) The geographic unit formed by combining all ZIP codes with the same three initial digits contains more than 20,000 people; and (2) The initial three digits of a ZIP code for all such geographic units containing 20,000 or fewer people is changed to 000;
- All elements of dates (except year) for dates directly related to an individual, including birth date, admission date, discharge date, date of death; and all ages over 89 and all elements of dates (including year) indicative of such age, except that such ages and elements may be aggregated into a single category of age 90 or older;
- Phone numbers;
- Fax numbers;
- Electronic mail addresses;
- Social Security numbers;
- Medical record numbers;
- Health plan beneficiary numbers;
- Account numbers;
- Certificate/license numbers;
- Vehicle identifiers and serial numbers, including license plate numbers;
- Device identifiers and serial numbers;
- Web Universal Resource Locators (URLs);
- Internet Protocol (IP) address numbers;
- Biometric identifiers, including finger and voice prints;
- Full face photographic images and any comparable images; and
- Any other unique identifying number, characteristic, or code (note this does not mean the unique code assigned by the investigator to code the data).
You should also be careful of deriving codes from identifiable information, such as using someone's initials, which could lead to discerning their name.
Most people are familiar with indirect identifiers such as demographics, which can include gender, age, sexual orientation, religious affiliation, race or ethnicity, etc. But other indirect identifiers could be the size of a person’s hometown and the nature of their community (e.g., urban, agricultural, suburban, etc.). A person's relationships and family structure also provide potentially identifying information: household size; number, ages, and gender identity of children; and domestic partnership status or changes. Even details of someone’s personal characteristics or hobbies could add to potential identification.
A combination of such indirect identifiers could still reveal someone’s identity even without direct identifiers. Especially with small sample sizes or fairly homogeneous populations, any variances might inadvertently point out a single person.
Example 1
A specific college course has 25 students: 12 identify as women assigned female at birth, 8 identify as men assigned male at birth, 3 identify as non-binary, and 2 identify as transgender (one woman and one man). In terms of age, 21 are traditional college ages (18-22), and 4 are non-traditional ranging from 29 to 51. For race, 14 identify as White, 3 as Black, 3 as Latinx, 2 as East Asian, 1 as Pacific Islander, and 2 as mixed race.
If a researcher surveyed this group on dating habits and sexuality and asked for gender, race, and age, someone looking at the data could reasonably link responses with certain students. This survey would not be anonymous, and depending on the nature of the questions, inadvertently revealed responses could lead to embarrassment or possibly even reveal someone's sexual orientation who was not publicly out.
Example 2
A researcher studying the gender gap in the finance industry wishes to interview finance professionals in three contiguous counties without big cities about job satisfaction and advancement. The researcher will not ask for individual names or names of institutions but will record gender, race, age range, position, and years in the profession, as well as the size of the financial institution. In such a primarily white, male industry, the combination of those demographics with the institution size could reveal certain individuals, which poses an economic risk if responses include criticism of the workplace.
Not at all. Studies that wish to compare data between different demographics or that seek to collect longitudinal data often need to collect extensive identifiers. But be intentional about the identifying information you collect: do not collect more identifiers (direct or indirect) than you reasonably need for your study.
- Any data you collect in person (e.g., interviews, video recording, etc.) can never be anonymous.
- Focus groups by nature cannot be anonymous, except, for example, a Zoom focus group in which participants keep their cameras off and use pseudonyms (provided participants wouldn't recognize one another's voices).
- If you link someone’s name, email, or phone number with responses, those data are not anonymous, even if you separate the identifiers from responses and use a linking code.
- Data are not anonymous if a combination of indirect identifiers (e.g., gender, race, age, etc.) could reasonably identify or reidentify a specific individual. Especially with small sample sizes or populations in which certain demographics are less common, you need to be careful about calling a study “anonymous.”
Confidentiality
Confidentiality in research refers to how researchers handle data. It represents an agreement (through oral or written informed consent) between you and potential participants that you will never disclose their individual responses and identities beyond the research team unless a participant gives you permission (preferably in writing) to use identifying data.
You can't guarantee absolute confidentiality, which you need to tell participants. For example, you can’t control whether members of a focus group share others' information. Or, in the case of a participant complaint, the HRPP or IRB might need to review data and possibly consult with appropriate WWU officials (e.g., University legal counsel). You also need to comply with applicable mandatory reporting laws, such as if a participant expresses the intent to harm self or others.
If you collect or transmit data online (e.g., via online surveys, Zoom interviews, email, Dropbox, etc.), a low risk of a confidentiality breach exists. For highly sensitive information, you should minimize this risk by encrypting data transmission.
Using the HRPP consent templates can help you include appropriate language about confidentiality in your consent materials.
Establish an effective data management plan and store all data securely to prevent loss or theft.
Avoid storing data on personal devices, even if protected by password or biometric authorization, because if you lose the device, an unauthorized person might gain access to identifiable data and breach confidentiality assurances. If you use your device to collect data (for example, recording interviews on your phone), be sure to transfer that data to secure storage as soon as possible.
In addition, WWU HRPP or other administration (or a faculty adviser, if applicable) might need to access your raw data in certain circumstances.
For these reasons, the WWU HRPP strongly advises storing all research data on WWU-supported cloud drives such as OneDrive or Teams.
Privacy
Privacy refers to an individual’s control over the extent, timing, and circumstances of sharing personal information (physical, behavioral, or intellectual). Privacy refers to people whereas confidentiality refers to data. Privacy is a right that can be violated whereas confidentiality is an agreement that can be broken.
Expectation of privacy means the reasonable belief that in a certain setting, people can speak and act freely without anyone observing or recording their personal information. Someone having a loud conversation on the bus or in a restaurant would not reasonably have an expectation of privacy. But elsewhere, people would not expect someone to be taking notes or recording for research purposes: classrooms, church services, private events, non-public online communities that require joining as a member, etc.
An event such as a wedding reception might involve a large crowd of people with individuals taking photos and videos on their phones. This scenario might seem public, but attendees would not expect a researcher to be observing and recording their behavior or words. On the other hand, if the reception took place at a public beachfront with nearby crowds, the reasonable expectation of privacy diminishes. (Although a researcher should likely refrain from collecting data out of respect for attendees.)
You should protect participants’ privacy during study recruitment, consenting, and data collection.
If you enter a collective space such as a classroom and explain your study, then ask for interested people to sign up on the spot, some individuals could feel uncomfortable or hesitant if your study relates to highly personal behaviors or specific beliefs. Similarly, if you interview people about sensitive topics, participants should feel comfortable in the data collection location—a private library study room might be more appropriate than a coffee shop.
Be mindful of settings and obtain any necessary permissions to conduct research activities in selected spaces.
Online spaces also require privacy considerations. For example, you would not need to obtain permission to collect data from an open, public social media page, where people understand anyone may see their posts and comments. But if an online space requires permission to join or is password-protected, you may not collect data without express permission of the administrator(s) and without notifying the community of your intentions as a researcher. You also need an administrator's permission to post recruitment information in a private online space.
With an increasing shift to Zoom, Team, or other online data collection, you also need to be respectful of the privacy of non-participants. If study participants decide to take part in an interview or focus group while in a public location or household with other individuals, you should ask them to use a blurred or virtual background to protect the privacy of non-consented individuals.