Privacy by Design for Engineers

SecurityPrivacyEngineeringCompliance
Share on LinkedIn Share on X Share on Reddit Share on HN Share on Bluesky

Legal asked for GDPR compliance six weeks before launch. Engineering had already logged full IP addresses, device fingerprints, and page titles in a single analytics_events JSON blob with no TTL. Privacy by design would have asked before the schema shipped: what do we need, for how long, who accesses it, and how do we delete it when users ask?

Privacy by design (Ann Cavoukian's framework) isn't legal paperwork — it's architectural decisions engineers make daily.

Seven principles translated for builders

  1. Proactive not reactive — threat model privacy in design docs
  2. Privacy as default — opt-in for sharing, not opt-out
  3. Full functionality — privacy without breaking core UX
  4. End-to-end security — encrypt in transit and at rest
  5. Visibility and transparency — users know what's collected
  6. Respect for user privacy — deletion actually works
  7. Accountability — audit logs, DPA records

Data minimization at the schema layer

-- BAD: collect everything
CREATE TABLE users (
  email TEXT,
  full_name TEXT,
  ip_address INET,
  device_fingerprint TEXT,
  browsing_history JSONB
);

-- BETTER: purpose-limited
CREATE TABLE users (
  id UUID PRIMARY KEY,
  email TEXT NOT NULL,
  display_name TEXT,
  created_at TIMESTAMPTZ DEFAULT now()
);
-- Analytics: separate table, pseudonymous user_id, 90-day TTL

Challenge every field: What decision does this enable? What breaks without it?

Remove:

Purpose limitation and retention

Tag tables with retention policy in code or metadata:

# data_retention.py
RETENTION_POLICIES = {
    "analytics_events": timedelta(days=90),
    "audit_logs": timedelta(days=2555),  # 7 years regulatory
    "session_tokens": timedelta(days=30),
}

Scheduled job deletes or archives — not manual quarterly scripts someone forgets.

DELETE FROM analytics_events
WHERE created_at < now() - interval '90 days';
-- Batch in chunks to avoid lock storms

Document purposes in privacy policy mapping to tables — legal maintains mapping; engineering implements TTL jobs.

Pseudonymization patterns

def pseudonymize(user_id: str) -> str:
    return hmac.new(PSEUDO_KEY, user_id.encode(), "sha256").hexdigest()[:16]

Store pseudonym in analytics; map key in separate vault with stricter access. Re-identification requires vault access — logged and rare.

Separate analytics database from PII database — analyst queries never join email without approval workflow.

Access control and audit

# Generate fake data for staging
pg_dump prod --schema-only | psql staging
python scripts/seed_synthetic_users.py

Prod data in staging is a privacy incident waiting for a laptop theft.

User rights: export and deletion

Engineering must implement:

Export (DSAR): aggregate user data across services — start with internal user ID index.

def export_user_data(user_id: str) -> dict:
    return {
        "profile": profile_repo.get(user_id),
        "orders": orders_repo.list_for_user(user_id),
        "support_tickets": tickets_repo.list_for_user(user_id),
    }

Deletion: cascade or anonymize — hard delete vs soft delete policy per regulation.

def delete_user(user_id: str):
    orders.anonymize_user(user_id)  # legal hold may block
    profile.hard_delete(user_id)
    cache.invalidate(user_id)
    search_index.remove(user_id)
    enqueue_third_party_deletion(user_id)  # Stripe, SendGrid, etc.

Deletion must reach backups on defined schedule — document backup retention vs erasure requests.

Privacy review in design process

Checklist for RFCs touching user data:

15-minute privacy pass in design review beats quarter-end scramble.

DPIA triggers

Trigger Data Protection Impact Assessment when launching features that process new categories of personal data, large-scale profiling, or systematic monitoring. Engineering RFC checkbox: "DPIA required?" routes to legal early — not two weeks before launch.

Operational notes

Contract third-party SDKs with same retention limits as first-party data — analytics SDK buffering events locally may violate deletion requests if vendor buffer TTL exceeds your policy.

Log data access to PII tables with purpose code — supports GDPR accountability principle and speeds breach impact assessment when logs show who touched affected records.

Review vendor DPAs when adding analytics SDKs — subprocessors list must match actual network calls from mobile builds verified in Charles or mitmproxy during QA.

Privacy engineering checklist

Privacy is not legal-only — engineers implement retention cron and deletion cascades.

Data minimization in API design

Request only fields product needs — optional profile fields become mandatory once form exists. Default API responses exclude sensitive attributes.

Privacy review gate in SDLC

Launch checklist: data categories, retention, legal basis, third-party processors, deletion path. Block launch without DPO sign-off on new PII surface.

Pseudonymization at ingestion

Hash user identifier with rotating salt for analytics pipeline. Engineering implements salt rotation without reprocessing historical events where possible.

Engineering metrics

Track count of services storing PII, average retention days, deletion request SLA compliance — platform dashboard visible to privacy team.

Threat modeling for data flows

STRIDE on features collecting new data — spoofing, tampering, repudiation, information disclosure, denial of service, elevation. Privacy-specific: linkability, identifiability, detectability. Output drives minimization and retention choices before sprint commitment.

Privacy budget for analytics

Cap distinct event properties containing quasi-identifiers per session — exceeding budget drops fields server-side even if client sends them. Engineering enforcement beats policy PDF alone.

Data classification tags in schema

Column comment or metadata tag pii:email consumed by linter blocking SELECT * into logs. Engineering-generated catalog from migrations feeds DPIA automation — new pii column opens Jira privacy subtask automatically in mature orgs.

Privacy regression tests

Automated crawl asserts no new third-party cookie without consent category — diff against baseline after marketing PR. Fails CI when doubleclick.net appears in network log of staging homepage without CMP load preceding.

Layered privacy defaults in feature flags

LaunchDarkly flag privacy_safe_mode defaults true in EEA geolocation — reduced analytics collection without separate build. Flag evaluation server-side; client cannot override geo-based defaults by manipulating local flag cache.

Closing notes

Privacy review checklist attached to Jira epic template — engineers cannot mark epic done until DPIA ticket closed or waived by DPO with documented rationale for low-risk internal tool.

Additional guidance

Engineering managers review epic closure checklist including privacy ticket — cultural reinforcement beyond automated linter. New hire onboarding includes thirty-minute privacy for engineers module explaining data classification tags used in schema comments and how to request DPO review when product brief introduces new personal data category such as biometric or precise geolocation beyond coarse IP derived region.

Data protection impact assessment links to architecture diagram stored in git docs/privacy/ — versioned alongside code changes affecting flows; auditor receives commit SHA pointing diagram matching production deploy not outdated Confluence export from prior year review cycle undermining trust in engineering privacy posture claims during enterprise sales security questionnaire.

Attach data-flow diagram to epic when collecting new PII category — DPO review gate in Jira blocks release until diagram merged in docs/privacy/ at same commit SHA as production deploy.

New analytics event properties require privacy reviewer approval in PR template checkbox — blocks merge when email or precise location added without documented lawful basis in adjacent comment.

Quarterly privacy office hours with engineering demos new CMP integration — attendance counts toward manager goal for privacy literacy; reduces last-minute launch blocks when legal discovers analytics SDK initialized before consent because engineer unaware of ordering requirement in frontend bootstrap sequence.

Resources

Frequently asked questions

What is privacy by design for software engineers?

Building privacy into architecture from the start — collecting only necessary data, limiting access, defining retention, enabling deletion, and encrypting sensitive fields — rather than bolting compliance on after launch. Proactive, not reactive.

What is the difference between anonymization and pseudonymization?

Pseudonymization replaces identifiers with reversible tokens keyed separately — still personal data under GDPR if re-identification is possible. Anonymization removes re-identification risk irreversibly. Engineering usually implements pseudonymization first; true anonymization is harder to prove.

Do engineers need legal approval for every data field collected?

Engineers should document purpose per field and challenge unnecessary collection in design review. Legal/privacy teams define requirements; engineers implement technical controls. Collaboration at schema design time prevents 'collect everything' defaults.

Hiring a senior Android / Flutter engineer?

I architect and ship production mobile software — Kotlin, Jetpack Compose, Flutter — for robotics, EV infrastructure, fintech, and real-time systems. Open to remote roles in Europe and the US.

Get in touch →