
Most guidance on this question implies that masking personal data equals compliance. It doesn't. Here is who owns which obligation, what redaction actually buys you, and the ten questions to answer before your next feedback upload.
Establish the controller and processor split in writing, mask personal data in preprocessing before the first analysis job, upload only the columns you need, and confirm retention and deletion actually execute including in backups. Pseudonymized data is still personal data under GDPR, so redaction is a safeguard rather than an exemption. Thematic operates as a GDPR processor, masks personal data before analysis as a best-effort priced add-on, and destroys customer data within 30 days of contract end to the NIST 800-88 standard.
Open-text feedback is the messiest personal data most companies hold. A customer typing into a survey box doesn't stay on topic. They give their name, their account number, the branch they visited, the agent who helped them, and sometimes a health condition or a complaint about a family member. Then that column gets exported and handed to an analytics vendor, and someone in legal asks whether that was allowed.
You analyze it lawfully by getting four things right, in this order: establish that you're the controller and your vendor is the processor, sign the contract that says so, mask personal data before analysis rather than after, and make sure retention and deletion actually execute. Thematic operates as a processor under the General Data Protection Regulation (GDPR). The customer decides what to upload, from whom, and why. Thematic processes it only on those instructions. Thematic can also mask emails, names, phone numbers, addresses, and identifiers before analysis begins.
One thing to settle first, because most guidance on this question gets it wrong: masking personal data doesn't take your feedback out of GDPR scope. Below is what redaction actually buys you, who owns which obligation, and the checklist to run before the next upload. This is an explanation of mechanics, not legal advice.
Pseudonymized data is still personal data under GDPR. European Data Protection Board guidance is explicit about it: pseudonymized data that could be attributed to a person using additional information remains information about an identifiable person. Article 32 names pseudonymization as an appropriate safeguard, and a safeguard lowers your risk. It doesn't end your obligation.
Two consequences that catch teams out:
Redaction reduces how much personal data reaches the analysis, and it shrinks the blast radius if something goes wrong. It doesn't convert a regulated dataset into an unregulated one. It doesn't remove your duty to define a lawful basis, honor data subject rights, or delete data on schedule.
The practical consequence: treat redaction as one control among several, not as the compliance step.
Most confusion on this question comes from teams assuming the vendor carries duties that are legally theirs. Under GDPR, the company collecting the feedback is the controller and the analytics platform is the processor. The controller decides the why and the what. The processor acts on instructions.
| Obligation | Who owns it | What that means in practice |
|---|---|---|
| Lawful basis for processing | You, the controller | Decide and document why you may analyze this feedback before you upload it |
| What data gets uploaded | You, the controller | The vendor can't decide your scope. If you upload a column you didn't need, that's your decision |
| Written processing contract | Both, required of you | Article 28 requires a written agreement. Get it signed before the first upload, not at renewal |
| Technical safeguards | Your processor | Encryption, access control, masking, isolation, and destruction to a stated standard |
| Answering a data subject request | You, with processor support | You respond. The vendor has to give you the means to find and delete the record |
| Retention and deletion | Both | You set the period. The vendor has to actually execute it, including in backups |
| Subprocessor changes | Your processor notifies you | You need to know who else touches the data, including translation services |
Four mechanics decide whether an analysis setup is defensible. None of them is exotic, and all four are usually skipped.
Mask before analysis, not after. Order matters more than teams expect. If personal data is stripped only at the reporting layer, it was already ingested, modeled, and stored in the clear. Redaction has to run in preprocessing, before the first analysis job, so the models never see the raw identifiers. Summaries and derived scores are built from the comment text, so anything that reads that text reads whatever you failed to mask.
Only upload the columns you need. The cheapest privacy control is not sending the data at all. Customer ID, email, and full name usually travel along out of habit rather than necessity. If you only need a pseudonymous key to join results back, send only that.
Know where non-English text goes. Multilingual feedback is often translated before analysis, which routes the text through a translation service. That service is a subprocessor in your processing chain, and your privacy reviewer will ask about it.
Make deletion provable. A retention policy that nobody executes is worse than no policy, because it documents an intention you're failing to meet. Ask for the destruction standard and the timeline in writing, and ask specifically what happens to backups, which are the usual exception.
Most of these are configuration choices in a feedback analytics rollout rather than anything a vendor decided for you.
The failures cluster:
One question to put to a vendor at demo time: ask them to show you a comment before and after redaction, on your own data, and then ask what their redaction misses. A vendor who names the failure modes has tested it. A vendor who says it catches everything hasn't.
Thematic is a processor, and the split is explicit. The customer solely determines what data to upload, from whom, and for what purpose. Thematic doesn't classify or represent customer data, and processes it only on the customer's instructions. Thematic signs a contract governing the processing of EU personal data. The customer keeps responsibility for data outside that scope, such as anything downloaded locally or printed.
Redaction runs in preprocessing, before analysis. Thematic masks email addresses, web addresses, unique identifiers, encoded data, physical addresses, names, and phone or number sequences, replacing each with a token such as [EMAIL] or [NAME]. It also strips quoted email history by detecting the "On ... wrote:" line and removing everything below it. Because summaries and synthetic scores are built from comment text, Thematic's guidance is to run redaction before the initial analysis job rather than after.
Redaction is best-effort and priced as an add-on. It isn't guaranteed to catch every instance, and it's configured through your customer success manager rather than switched on by default. That's worth knowing during evaluation rather than discovering later.
Dataset-level modify and delete supports erasure. Customers can modify and delete individual datasets, which is what makes an erasure request actionable against analyzed data.
Retention and destruction are stated, including the backup exception. Customer data is kept for the contract lifetime and deleted within 30 days of termination using NIST 800-88 sanitization. Backups are the exception and are purged on their own rotation. Personal data on a user account is kept for the account duration plus three months. All customer data is classified as Restricted, the most protective of the three tiers Thematic uses.
Subprocessors and transport are governed. Customers are notified of subprocessor list changes under a documented policy. Customer data moves only through approved channels, transfer by USB drive or email is prohibited, and access is from encrypted laptops only. Non-English feedback is translated into English via Google Translate so themes stay consistent across languages, with the untranslated source preserved alongside it.
Thematic's full posture is published on its security and compliance page. For the certification and audit side of a vendor review, including SOC 2 scope and penetration testing, that's a separate question with a separate answer. So is whether the EU AI Act's emotion-recognition ban touches sentiment analysis.
Ten questions, aimed at your own team as much as at the vendor. They sit alongside the wider requirements worth putting in an analytics RFP.
Establish the controller and processor split in writing, mask personal data in preprocessing before the first analysis job, upload only the columns you need, and confirm that retention and deletion actually execute including in backups. Thematic operates as a GDPR processor. It masks personal data before analysis as a best-effort priced add-on. It supports erasure through dataset-level modify and delete. It destroys customer data within 30 days of contract end to the NIST 800-88 standard.
Do not accept redaction as your compliance story. It's a safeguard that lowers exposure, and the obligations sit with you either way.
Thematic turns fragmented feedback into one consistent source of customer truth — so every team acts on the same customer story. Up and running in days, not quarters.

Transforming customer feedback with AI holds immense potential, but many organizations stumble into unexpected challenges.