What Your Compliance System Needs to Track
Data privacy compliance is often focused on outputs: privacy notices, data processing agreements, cookie consent banners, and the like. It's common for smaller orgs to do a big push to get information together to put these public artifacts in place, without much thought to how they'll be maintained in the long run.
But compliance is a long game, and keeping pace with your obligations—including what goes in those public outputs—requires regular updates. If every update requires ad hoc check-ins with the various teams responsible for your systems and data flows, compliance will be costly and error-prone.
Running a mature compliance program means centralizing the key inputs and implementing processes to refresh them as your practices and partnerships evolve. That way, you can regularly update your outputs by reference to a single, central repository of data, and keep your ongoing compliance costs low.
The system that's right for your organization will depend on your size, systems, and existing processes. Options range from a handful of spreadsheets to commercial compliance platforms of different capabilities and costs. Whatever system you use, it should be customized to your actual needs. Tracking data you don't use is a recipe for busywork and frustration, and for letting your "source of truth" rot for lack of updates. If you're tracking data that doesn't make its way into any of your outputs, prune it.
Your compliance obligations will determine your outputs, and therefore what you need to track. The framework descrived below accounts for requirements specific to GDPR, GDPR's cross-border transfer regime, and various state laws. Before diving in, it might be helpful to read our previous posts to get a sense of which of these apply to you:
- Does GDPR Apply to Your US-Based Service?
- Are You Making Any Cross-Border Data Transfers Under GDPR?
- Twenty-Five State Privacy Laws: Thresholds-and-Exemptions Matrix
The questions your system needs to be able to answer
Here are the questions that a mature data privacy compliance program needs to be able to answer confidently:
1. Where does personal data live, and who has access?
Your system should track every datastore or system holding personal data, who is responsible (internally or externally), and what outside parties have access to it—whether as an attribute of the system (e.g. the public, registered users), or to support its operations (e.g. vendors via an API).
This information is the foundation of several compliance artifacts and the starting point for just about every compliance response/process, including:
- Records of processing activities (RoPAs) required by GDPR Art. 30
- Data breach response, as dictated by GDPR and every state data privacy law
- Data subject request (DSR) fulfillment
- Retention enforcement under GDPR Art. 5(1)(e) and various other laws
- Security program scoping under GDPR Art. 32 and any other law requiring security measures "appropriate" to your processing
- Risk assessments and audits, including under GDPR Art. 35 and any law requiring evaluation of processing risk
2. What do you do with data, and on what basis?
Track everything you do with personal data, including the purpose of the processing, the legal basis, the categories of data involved and who it pertains to, and who receives it. If GDPR compliance is a goal, then you should record a legitimate interest assessment (LIA) for any processing done under the "legitimate interest" basis.
This information is used in the following artifacts and processes:
- Records of processing activities (RoPAs) required by GDPR Art. 30
- Privacy notices
- Legitimate interest assessments when the GDPR legal basis is legitimate interest
- Purpose-limitation checks when your organization wants to use existing personal data for new purposes
3. Which data is sensitive?
GDPR and many US laws require careful handling of "sensitive" personal data, which variously includes race and national origin, political opinions, religious beliefs, sexual orientation, union membership, genetic and biometric data, health-related data, precise geolocation data, and data about children, among other categories.
Your system should enable you to identify any sensitive data you process. Besides tracking any processing categories as you must for all personal data (see the first question), you may have special consent requirements for sensitive data, perform risk assessments regarding its use, and avoid sharing it under defined circumstances.
This information is used in the following artifacts and processes:
- Data privacy impact assessments and other risk assessments required by GDPR Art. 35
- Privacy notices
- Consent architecture
4. Whose data do you hold, and what's the process for correcting or deleting it?
If you receive a data subject request to provide, correct, or delete data about a specific person, you need to be able to quickly find all of the data about that person, and update or delete it as the law may require.
This information is used in the following time-sensitive processes:
- Data breach response
- Data subject request (DSR) fulfillment
5. Who else gets the data?
You must be able to determine easily which third-parties have access to personal data that you hold. This includes the legal/contracting entity, their location, their GDPR-defined role (e.g. controller, processor, joint controller), information about your contract with them, and a reference to their subprocessor list.
This information is used in the following artifacts and processes:
- Records of processing activities (RoPAs)
- Cross-border transfer analysis under GDPR Chapter V
- Transfer impact assessments (TIAs) which must be produced on demand under GDPR standard contractual clauses (SCCs)
- State law sale/sharing analysis
6. Where does data cross borders, and under what mechanism?
If you're subject to GDPR, then cross-border transfers of data to non-EEA countries must be subject to adequate safeguards, which are covered in Are You Making Any Cross-Border Data Transfers Under GDPR?
For any flow of personal data between you and a third party, you must be able to determine whether it's a "third-country transfer" under GDPR, which direction the data flows (i.e. who is the "importer" and who is the "exporter" under GDPR Chapter V), and how you're ensuring that transfers are adequately protected (e.g. SCCs, compliance with the Data Privacy Framework, or another mechanism). If you're subject to GDPR, you must record a transfer impact assessment (TIA) for third-country transfers relying on safeguards like the SCCs (transfers relying on adequacy decisions, including the DPF, don't require one).
This information is used in the following artifacts and processes:
- Records of processing activities (RoPAs)
- Cross-border transfer analysis under GDPR Chapter V
- Transfer impact assessments (TIAs) which must be produced on demand under GDPR standard contractual clauses (SCCs)
- DPF certification and recertification
7. What did data subjects specifically agree to (and what did they decline)?
Consent is one of the legal bases for processing personal data under GDPR, but it comes with unique limitations: a data subject must be able to withdraw consent at any time. If they do, you need to be able to stop any processing that they withdraw their consent for, if that's your only legal basis for the processing.
In addition, many US state laws prohibit the processing of "sensitive" personal data (variously defined) without the data subject's express consent. For both of these reasons, you need to track when and how data subjects grant consent to process their personal data and for what purposes—and when they revoke their consent.
Finally, an increasing number of state laws require you to enable data subjects to opt out of data sales and target advertising—and some specifically require online platforms to honor global opt-out signals, such as "Do Not Track" features in phones and browsers. So it's as important to track when and under what circumstances users have specifically refused to permit the processing of their personal data.
This information is used in the following artifacts and processes:
- Data subject request (DSR) compliance
- Consent architecture
8. How long do you keep it, and can you prove you deleted it?
Storage limitation—closely related to data minimization—is a fundamental principle of the GDPR and a requirement of most data privacy laws: you can only keep data as long as you reasonably need it for a legitimate purpose. This means that all personal data must have a retention period or rule associated with it that is scoped in a justifiable way, and the retention rule must be implemented and function in practice. If you become subject to a legal hold (e.g. if you are threatened with relevant litigation or receive a government subpoena), you must also be prepared to preserve data that would otherwise be automatically deleted.
This information is used in the following artifacts and processes:
- Retention policy compliance
- Privacy notices
9. What has gone wrong, and who has come asking?
GDPR Art. 33(5) requires you to document all data breaches, including those that don't require notification to supervisory authorities (because they are "unlikely to result in a risk to the rights and freedoms of natural persons"). This documentation must record "the facts relating to the personal data breach, its effects and the remedial action taken."
You should also keep track of any government or law-enforcement demands for personal data, and your response to them. The absence of requests (or your refusal of them) may support a risk-based TIA conclusion, but you should be ready to provide documentation if a data protection authority comes knocking.
This information is used in the following artifacts and processes:
- Data breach notification compliance
- Transfer impact assessments (TIAs)
10. What did you decide, and when did you last check?
GDPR and some US data privacy laws require you to maintain internal documentation of various determinations you make about the processing of personal data, and other compliance measures, including:
- Data privacy impact assessments, required by GDPR Art. 35 to justify high-risk processing of personal data, such as processing using "new technologies," as well as similar assessments required by some state laws.
- Transfer impact assessments, required under GDPR when transferring data to "third countries" whose laws don't protect personal data to the same extent as GDPR.
- Legitimate interest assessments, required by GDPR Art. 6 when relying on your "legitimate interests" to process personal data, to ensure those interests aren't "overridden by the interests or fundamental rights and freedoms of the data subject."
These determinations must be updated when the details of the processing (including what is processed, the purposes, or the legal circumstances) change.
Your program should also track and keep updated any policies related to data protection, record training materials and compliance with training requirements, and regular reviews of your compliance materials.
In addition to the processes listed above, this information is used in the following artifacts and processes:
- Compliance certifications required for DPF certification and for certain businesses under California's CPPA regulations.
- Customer procurement diligence, an increasingly common fact of life for companies with EU customers.
Compliance system design guidelines
That's the information your compliance system needs to track, at a high level. Implementation will depend on the specifics of your organization and its existing systems, but here are some common-sense design constraints that any data privacy compliance system would do well to implement:
- Don't repeat yourself. Every fact that your system tracks should live in a single place, whether that's a spreadsheet, database table, or compliance platform. Any other register that references that fact should link to its home base.
- Traversable links. Links between registers containing different information should be traversible in both directions. For example, if you have a list of systems containing personal data, and those systems are referenced in your register of data flows, the data-flow register should link back to entries in the systems list. From the systems list, it should be possible to get a list of all data flows involving that system.
- Ability to derive required views. Various compliance outputs and processes require different views on the data described above. Your system should be designed to pull these views together from the data, rather than requiring the data to be stored in a view-specific way. For example, records of processing activities required by GDPR Art. 30 would pull from records describing third-party recipients, personal data categories, data flows, and retention periods.
- Event-driven updates. The records and outputs of your system will require regular updating. These updates should be driven by events, not by the calendar. For example, when you add a new vendor that will get access to personal data, that should drive updates to your records of: the systems containing personal data, recipients of that data, data flows, and (potentially) cross-border transfers. If these updates wait for an annual whole-system review, you're likely to miss things, or publish inaccurate outputs between reviews. A register maintained by calendar tells you what someone remembered; a register maintained by events tells you what happened.
- Tune the level of detail to the demands of your consumer. The level of detail you track about the items above will depend on how you need to use it. If you're implementing an automated system for compliance with data subject access requests, it makes sense to track personal data at the level of individual data fields. On the other hand, if DSRs will be processed manually, then the highest level of detail you need for any output might be the category level, e.g. as required for GDPR Art. 30 records of processing activities. When in doubt, it's safer to track data at a finer degree of resolution than a coarser one.
Any number of system architectures can meet these requirements, whether you're relying on spreadsheets, bespoke internal databases, or compliance platforms—and regardless of whether your data architecture is centered around your systems, data elements, or the object model of your favorite compliance platform.
The minimum viable dataset for US-only compliance.
If GDPR is not a concern for you (if you're not sure, see this post), then the bar is lower. A system capable of answering questions 1, 3, 5, and 7 above will suffice when coupled with a log of data requests (by both data subjects and governments/law enforcement). The system can grow organically to answer other questions as your compliance and customer needs demand it.
Take action.
This week, evaluate your current systems against the ten questions—for each question, how long would it take you to answer it? Build a roadmap based on the gaps, prioritizing the artifacts you'll need the soonest.