Field Notes

Every team closed its ticket but the customer's data is still there.

The security problems that matter at a growing SaaS company cross every team. Nobody owns them end to end. Fix that.
September 29, 2026

A customer asks you to delete their data. Support files a ticket. The app team deletes the rows in the main database and closes it. Done.

Except the data team's nightly job was supposed to catch the soft-delete flag, and it filters on updated_at, which the flag flip didn't touch. The search index rebuilds weekly. Last quarter an engineer exported that customer's records to a CSV in a shared drive to debug a billing issue. The analytics SDK has been sending their email address as a user property since launch. Backups keep everything for 35 days.

Every one of those systems has an owner. Every owner did what their ticket asked. Nobody can tell the customer, with evidence, that their data is gone.

We see this pattern at growing SaaS companies, and it isn't limited to privacy. It's also how most real breaches work. An attacker's path starts in one team's system and ends in another's. A support tool holds customer session tokens. A CI job holds production credentials. A contractor left in March and their API key still works. Each piece has an owner, but the path between them doesn't.

This post is about who should own that path. We'll cover what a cross-domain problem owner does, why a ticket queue can't do the same job, and how a company of 30 to 200 people can set this up without hiring a department.

Attack paths cross the org chart

Most security programs are organized by component. One person owns the AWS org. Another owns Auth0. A team owns the API. When a finding comes in, it goes to whoever owns that component, and the ticket closes when they ship a fix.

That's fine for bugs that live in one place. It falls apart for the questions that decide whether a compromise becomes an incident. Those also happen to be the questions a serious buyer's security team asks on the follow-up call.

Can one customer read another customer's data? Check every path: the API, bulk exports, background jobs, caches, search, support tools, and any AI feature that does retrieval.

When someone leaves, is their access gone everywhere? The IdP is the easy part. The hard part is the GitHub personal access tokens, the CI secrets they created, the SSH keys on the bastion hosts if you're still using those, the IAM user someone made outside SSO, and the database grant from that incident last year.

When you say support access is scoped and time-limited, is that enforced in every service, or just in the admin panel?

Where can customer content leave your main environment? Vendors, SDKs, logs, exports, LLM providers.

If an AI agent in your product must not reach a resource, is that enforced by its credentials, or just requested in its system prompt?

Each of these is a claim about the whole system. No single team can answer it. You answer it the way an attacker would: list every path to the thing you care about, try each one, and see which ones work.

Two public incidents show what this looks like. In October 2023, Okta disclosed that an attacker used a service account's credentials to get into its support case system. From there they downloaded HAR files that customers had uploaded for troubleshooting. Some of those files held valid session tokens. In January 2023, CircleCI disclosed that malware on one engineer's laptop stole a session token that had already passed SSO and 2FA. That token reached production systems holding customer secrets. Next year someone's going to disclose that an infostealer grabbed ChatGPT history logs which had some other critical credential material leading to some egregrious breach.

Look at who owned what in each case. The laptop, the identity layer, the support tool, and the customer data all had different owners. The attacker walked through all of them.

Why it gets worse after a few dozen people

At a dozen people, a few engineers hold the whole system in their heads. They know who has production access. They know where customer data goes. They remember which shortcuts they took and why.

Somewhere past 30, that shared picture breaks up. Teams make design choices on their own. Each new integration adds a credential and a data flow. Analytics, support tools, and AI features find new uses for customer data. Sales signs deals that promise things the product can't do yet. And AI coding tools now write a big share of the code, faster than anyone can keep up with what it does.

The cheapest attack paths are the ones nobody remembers making. The staging admin endpoint that quietly made it to prod. The long-lived AWS key in a CI variable. The S3 bucket from a sales demo two years ago. None of these need a clever exploit - they just need someone to happen across them or go look.

Most companies at this stage do plenty of security work. There's a scanner, a compliance platform, a pen test, a backlog. But the work is sorted by component, and the exposure sits between components.

Write down what can't fail

Before you assign anyone, write down the few outcomes that would really hurt if they failed. Make each one a statement you could test. Categories like "data protection" don't count.

Here's a starting set most SaaS companies can adapt:

  1. A user in tenant A can't read or change tenant B's records by any supported path.
  2. Nobody can change production code or infrastructure except through reviewed CI, and CI credentials don't work from outside CI.
  3. Support staff can't act as a customer without a time-limited, logged grant scoped to that customer.
  4. Customer content only goes to the vendors on your DPA subprocessor list.
  5. A deletion request reaches every store in your processing inventory within the window you promised.

Rank by how easy it would be for an attacker to break each one today. If a property has a known easy path, move to the top. If breaking it would take three unknown bugs chained together, move it down.

No compliance framework will write this list for you. A SOC 2 audit checks that you run access reviews. It won't check whether your support tool can log in as any customer without leaving a trace.

Carry one problem all the way around

Security work splits into four activities that feed each other. You understand the system by knowing the business and theats, reading code, mapping data flows, and learning from incidents. You challenge your assumptions by trying to break what you think works. You build the changes that make a property withstand adversity. You operate the detection and response that tell you when those are pressured or break. What operations learns feeds the next round of understanding, and the loop starts over.

Big companies have a team for each activity. At 60 people, it's the CTO, two senior engineers, and maybe one security hire, all doing a bit of everything. The failure is the same at both sizes. Each activity happens, but rarely one person carries a single problem through all four.

Here's what that might look like. Say the question is whether your AI agent can reach a sensitive resource.

First, you map what the agent can reach. List every credential it holds and everything those credentials can touch, directly or through other tools. Include internal APIs that trust anything on the network instead of checking who's calling.

Then you try to break it. Maybe a different tool has a broader credential. Maybe an internal service trusts anything inside the VPC. Maybe a document the agent retrieves contains an injected instruction that tells it to call a tool for the attacker. That last one is a textbook confused deputy: the agent uses its own authority on someone else's behalf.

Then you fix the structure. Give the agent a token scoped to the user who's calling it, not a service account that can read the whole org. Replace broad tools with narrow ones, so query_database becomes get_invoice(org, invoice_id). Check authorization in the service the tool calls, not in the prompt. Require a human to confirm anything destructive or external.

Then you watch it. Log each tool call with the user, tenant, and arguments. Alert when a call gets denied for crossing a boundary.

Then you rerun the original attack and confirm the authorization layer blocks it. If the model just decided to refuse, that doesn't count.

That's five steps, at least three teams, and one outcome. I keep coming back to the difference between a request and a control. A system prompt that says "never access payroll data" is a request. A credential that can't read payroll data is a control. The request works when the model behaves. The control works when it doesn't.

Three roles per problem

You don't need a new department. You need a few named people for each problem.

The problem owner builds an end-to-end picture, finds gaps, coordinates the fixes, and gathers proof that they're closed. This is an assignment, not a job title. It could be a security engineer, a staff engineer, the platform lead, or someone from outside.

System owners change and run their own systems. The problem owner doesn't take over the data pipeline. The data team fixes the data pipeline.

The sponsor, usually the CTO, settles priority fights, frees up engineering time, and makes the call when a fix is expensive.

Drop any one and the work stalls. With no problem owner, every team finishes its piece and the outcome stays open. With no system owners, the problem owner ends up rewriting other teams' code. With no sponsor, the work loses to the roadmap every sprint.

The problem owner keeps one readable working doc. It says what the property is and why it matters, whether that's a contract clause, a legal duty, or a realistic attack path. It lists what's known, what's assumed, and what nobody has checked. It names the systems, teams, and vendors involved. It says what evidence exists today and what that evidence doesn't prove. It lists each needed change with an owner. And it says what "closed" means and who owns the property afterward.

That last part matters most. A property is closed when it's been shown to hold within a stated scope and someone owns keeping it that way. In practice that usually means a regression test in CI plus a signal in production. A sync meeting doesn't close it. A column of resolved Jira tickets doesn't close it. A clean scan doesn't either.

Example: deletion and access requests

This stops being so abstract with privacy issues, because the law spells out what outcomes you need.

Where GDPR applies, Articles 15 and 17 give people the right to see and erase their personal data. You generally have one month to respond, with extensions for complex cases. The European Data Protection Board's guidelines on the right of access (01/2022) push controllers to prepare for requests before they arrive. Processors have to help, but the controller owns the answer. Several US state laws, including California's CCPA as amended by the CPRA, set up similar rights with their own deadlines.

Take the sentence "we honor deletion requests" apart and you hit a cross-domain problem right away.

You need an inventory of what personal data you process, why, where it lives, who receives it, and how long you keep it. It has to match what you find in the code and the cloud, not just what's in a spreadsheet.

You need identity resolution. The same person is an email in the app database, a user_id in the warehouse, a hashed ID in the analytics tool, and a ticket requester in the support system. Without a mapping, good luck finding their data.

You need an access request process that verifies the requester, pulls data from every store, redacts other people's information, delivers the result safely, and logs all of it.

You need deletion that accounts for all this. A common design puts a deletion event on a queue. Each store has a consumer that deletes or anonymizes its copy and writes back an acknowledgment. A tracker marks the request complete only when every registered store has acknowledged. If one doesn't, someone gets paged.

You need retention that ideally runs on its own: TTLs, S3 lifecycle rules, or scheduled jobs, with logs that prove they ran.

You need a backup decision, likely made with counsel. Either you restore and delete, or you let backups expire on schedule and replay deletions after any restore. Write down which one you chose.

And you need to catch new things. When a feature adds new data, a new purpose, or a new vendor, that should trigger a review. When a new database, bucket, or SDK shows up that isn't in the deletion registry, someone should notice.

There's a security angle here that privacy write-ups tend to skip. The access request process is an attack surface. At Black Hat USA 2019, James Pavur presented results from sending GDPR access requests to about 150 companies while posing as his fiancée, with her consent. A meaningful share of them sent her personal data after weak identity checks or none at all. If your verification is "reply from the email on file," anyone who controls that inbox gets the data. Verify requesters as carefully as you verify logins.

Counsel decides what the law requires and signs off on exceptions, like backup handling or tax records. The problem owner figures out what the system does, what's feasible, and what evidence shows it works. An engineer shouldn't invent a retention exception to hide a gap. A lawyer shouldn't assume the deletion job reaches the warehouse.

Try this. Create a test account and use it normally for a week. Submit a deletion request through your real process. A week later, search for its email and user ID in your app logs, analytics tool, support system, warehouse, and object storage. If that takes an afternoon, you have a baseline. If you can't do it at all, that might be worth addressing.

Look for the problems nobody filed

A security team that works from a backlog fixes problems somebody already noticed. You need that. But it's not the whole picture. Attackers use the paths nobody wrote down, because those paths aren't in anyone's picture of the system.

So the problem owner needs protected time to go looking. The simplest way is a recurring investigation into one question. For example: where can customer content leave our main environment?

You answer it by comparing sources that ought to agree and often don't. Architecture docs and interviews tell you what people think happens. Code, dependency manifests, and config tell you what's built, including the analytics SDK from two years ago that ships request bodies to a third party. Your cloud inventory and IAM policies tell you what exists and what it can do; AWS Config and IAM Access Analyzer are cheap places to start. Egress logs, VPC flow logs, and DNS logs tell you what's observed, if you log them. Vendor contracts and your subprocessor list tell you what you've promised.

Every place those disagree is a work item. An analytics destination that isn't on your subprocessor list. A forgotten export job. A Lambda with s3:* on *. These often deserve more attention than anything in the vulnerability queue, and they're exactly what an attacker doing recon finds first.

The result doesn't have to be a vulnerability. It can be a corrected system map, new logging, or a newly named risk. It should always list what's still unknown, with an owner. Something like: "Deletion covers these five stores. We are confirming whether this vendor keeps derived records." Simple, with a next step.

One note for leadership. A quiet security queue might mean nobody is looking. A low finding count proves nothing on its own.

Example: testing AI-written code at the speed it's written

The second place this model pays off is AI-assisted development.

The security requirements don't care who wrote the code. What changes is volume, understanding, and correlated mistakes. There's a lot more code. The person merging it may understand it less. And AI-written tests tend to share the blind spots of AI-written code. A test that checks "a user can fetch their own invoice" passes whether or not the endpoint also hands out everyone else's. Nobody can read every generated line, and pretending otherwise just moves the bottleneck.

What works better is testing in several tiers.

The first tier runs on every pull request, is deterministic and has to be fast. Secret scanning with gitleaks or trufflehog, dependency scanning, targeted static analysis with Semgrep or CodeQL rules you've tuned, and your security regression tests. Only checks that are reliable and quiet should block a merge.

The second tier runs asynchronously against deployed builds. It lives outside CI's time budget, against a staging environment seeded with a few tenants. An agent looks at what changed, guesses how it could fail, runs dynamic tests, and tries to reproduce anything suspicious.

The third tier is people. It runs on a schedule and after large changes. Humans reassess the architecture, look for chained behavior, and poke at the parts nobody has touched in a while. This is where a tageted pen test usually fits, and where multi-step authorization bugs and business logic flaws usually show up.

The async tier stays tied to CI even though it runs outside it. A risk averse release rule is to hold changes to authentication, authorization, or tenant-scoping code until the async run finishes.

For the async tier to work, the agent needs some up front work. That means the diff and nearby code, the OpenAPI spec, deployment config, your authorization rules, logins for test users or API tokens in at least two tenants and two roles, and perhaps past findings. Then it follows a fixed sequence. It maps what the change touches, writes down specific ways it could fail, picks tools and test identities, runs the tests, and reproduces any hits. Each confirmed issue becomes a permanent regression test.

You don't have to build the execution layer from scratch. OWASP ZAP's Automation Framework runs YAML-defined plans with authentication, OpenAPI import, and scripting. Burp Suite supports authenticated scanning with recorded login sequences. Start with one scanner plus a small custom runner for multi-step authorization tests. Pytest, an HTTP client, and two-tenant fixtures are enough.

Scanners are good at finding generic bugs cheaply. The agent earns its keep by using what it knows about your system to aim tests at the specific, expensive bugs.

Say an AI-assisted PR adds async customer exports. A scanner might catch an injection bug. A targeted investigation asks the questions an attacker would:

  • Does the worker recheck authorization when the job runs, or only when it's queued?
  • If the requester loses their role while the job waits, does the export still finish?
  • Does the worker get the tenant ID from the job's authenticated context, or from a payload field the API copied from the request?
  • Is the download link a short-lived signed URL tied to the requester, or a guessable path?
  • Does the deletion registry know about the exports bucket?

You can't answer those with one request and one response. You need multi-step tests that watch state change over time. And look at the last question. It's the deletion problem again. These problems keep running into each other.

A few guidelines to keep things on the rails.

The agent is an attack surface too. Repo content and app responses are untrusted input, so prompt injection is a real risk. Run generated test code in an isolated container with staging-only credentials and limited network access. Don't let it edit the authorization spec or delete its own run records.

Before you trust any of this, measure it. Run it against old bugs from your own tracker, bugs you plant on purpose, and builds you know are clean. Track how many known bugs it finds, how often its findings reproduce, how many false positives each run produces, and what share of runs got past login.

Who should own it

The problem owner needs enough technical depth to trace a problem through code, infrastructure, and data. They need enough standing that other teams will change their systems when asked. And they need protected time, so the work doesn't vanish the first time a feature slips.

At a 30- to 200-person company, that's usually one of three people. It could be a dedicated security and privacy engineer whose main job is building protections, but who also helps test and operate them so the fix solves the real problem. It could be a staff or principal engineer who already chases problems across team lines and gets formal time and authority to do it. Or it could be an outside practitioner working as a fractional engineering lead, paired with an internal engineer who does the implementation and eventually takes over.

It helps a lot if the owner has seen these systems get attacked. Someone who has broken tenant isolation in other products will check the export worker, the cache key, and the search index. Someone who hasn't might check the API and stop.

Whoever it is needs real authority, or they end up blamed for problems they can't fix. At minimum, they need a seat early in the architecture and product decisions that matter, and read access to code, config, and telemetry. They need the right to set technical acceptance criteria within agreed policy, committed hours from system owners, and a clear path to escalate to the sponsor.

If the same person also handles every questionnaire, scanner alert, vendor review, and incident, you haven't created a problem owner. You've created a first security hire with a very long job description, and the cross-domain work will lose to interruptions most weeks.

Start simple

For most companies at this stage, tenant isolation or data deletion is the right first choice. Enterprise buyers ask about both. Attackers go after the first. Regulators care about the second.

Then run the loop on that property:

  1. Write it as a testable statement and record who approved it.
  2. Name the problem owner, the system owners, and the sponsor.
  3. Trace it through the real system. Compare the docs to the code, the cloud state, and what happens at runtime.
  4. Try to break it, starting with the cheapest attacker paths. Add agent-assisted testing where it helps.
  5. Fix what breaks, in the system that broke. Prefer structural fixes, like scoped credentials and server-side tenant lookup, over new process.
  6. Add a production signal so a failure gets noticed.
  7. Retest, and write down what's covered and what isn't.
  8. Name who owns it after closure, and put its regression test in CI.

Doing this once tests two things. It shows whether your discovery finds real gaps. It also shows whether your organization can carry a finding all the way to a lasting fix. In my experience, the second answer is the more surprising one.

After that, one question a month keeps you honest. Can you say what changed in the product or infrastructure, why it mattered, how you checked it, and who owns it now?

‍

Get Started

Let's Unblock Your Next Deal

Whether it's a questionnaire, a certification, or a pen test—we'll scope what you actually need.
Smiling man with blond hair wearing a blue shirt and dark blazer, with bookshelves in the background.
Noah Potti
Principal
Talk to us
Suspension bridge over a calm body of water with snow-covered mountains and a cloudy sky at dusk.