What is a Confluence PII and secret scanner?
A Confluence PII scanner checks documentation for patterns that may indicate personal information. A secret scanner looks for credentials, private keys, and token formats that could have been pasted into a page or uploaded with a text attachment. The two tasks belong together because working documents can contain both customer information and technical credentials.
Ravenshift’s PII & Secret Scanner for Confluence combines these checks with a review workflow. The current release build uses deterministic rules inside Atlassian Forge. It does not call a third-party AI or scanning service, and its manifest configures no external egress. This page describes the build being prepared for deployment; the public Marketplace release is coming soon.
A pattern match is a finding to investigate. It is not proof that the value is sensitive, and a scan without findings is not proof that the document contains no sensitive information. Reviewers need to consider the source, the scan coverage, and the context of each result.
Which Confluence content does it scan?
The app scans page-body text and bounded, supported text-like attachments. Supported extensions include TXT, JSON, CSV, TSV, LOG, ENV, YAML, XML, and Markdown. Attachments reported by Confluence with a text, JSON, or CSV media type can also be scanned within the same size limit.
| Content or operation | Current release limit or behavior |
|---|---|
| Page body | Up to 1,000,000 characters |
| Supported attachment | Up to 256 KiB per attachment |
| Stored findings | Up to 500 per page scan |
| Automatic changed-page scans | 1,000 pages per tenant per UTC day |
| Manual page and space scans | Separate 1,000-page daily budget |
| PDF, Office documents, images, encrypted files | Not parsed by this release |
Unsupported, oversized, failed, or skipped attachments produce incomplete coverage. Finding truncation and oversized page bodies also make coverage incomplete. Disabling attachment scanning does not make attachment content clean; it leaves coverage incomplete.
These distinctions matter when reviewing mixed documents. A page with a short paragraph and an attached PDF may have its page text checked while the PDF remains unexamined. Record that limitation and arrange another review for the attachment before treating the whole document as assessed.
Start with a representative test space
After the app becomes available, begin with a small space that represents your documentation patterns. Use synthetic examples rather than real customer records or working credentials. Include ordinary documentation, a synthetic email address, a clearly fake credential, and an unsupported attachment to check both findings and coverage warnings.
- Review the requested Confluence scopes and the release privacy disclosures.
- Confirm automatic changed-page checks and attachment scanning in Settings.
- Review the built-in detectors in What to look for.
- Open Governance Check on a test page and choose Check now.
- Review the result, redacted previews, and coverage information.
- Check the dashboard, change history, and exported evidence.
This acceptance exercise should answer practical questions: can the reviewer find the source page, distinguish a page finding from an attachment finding, and understand why a result needs attention? A useful rollout depends on that workflow as much as on the detector rules.
Review redacted findings in context
The app’s Findings / Issues view supports filtering by text query, space, severity, detector, and status. Findings use redacted previews rather than storing the original matched value. Reviewers can follow links to Confluence content they are allowed to see and investigate the context there.
Suppression marks an accepted false positive in app-owned state. Restore returns a suppressed finding to the open state. Neither action edits the Confluence page. If content needs correction, an authorized editor must update the source document through the normal content workflow.
Define who can accept false positives and what evidence they should record. A synthetic example in a training page may be intentional, while a similarly shaped value in a production runbook deserves investigation. Do not use suppression simply to reduce a dashboard count without resolving the underlying question.
The scanner does not automatically redact page content or rotate credentials. Route confirmed findings to the responsible content or security owner for remediation, then check the updated content again.
Use automatic checks and space reviews together
Changed-page and attachment events enqueue the affected page for a scheduled worker. This is bounded background processing, not an instantaneous guarantee that every edit has already been checked. Automatic work has a per-tenant daily budget.
Administrators can also request a space review. These reviews are resumable and use a separate manual daily budget. A large space can continue across worker invocations and, when necessary, across days. Plan the rollout around the saved scan state rather than assuming that scheduling a review means the whole space is finished.
The Spaces view distinguishes Not scanned, Up to date, and Needs attention. An unscanned space is not shown as Up to date. For a decision about a specific page, inspect its version and coverage details as well as the broader dashboard status.
Configure detectors carefully
The built-in rules cover common PII and credential patterns, including email addresses, phone numbers, payment or account identifiers, passwords, private keys, and provider-specific token formats. Administrators can enable or disable rules to fit their environment.
The current build also supports up to 50 custom regular-expression checks. Custom checks use case-insensitive/global matching and Medium severity; invalid or unsafe expressions are rejected. Start with narrow expressions and representative synthetic fixtures instead of attempting to match every possible identifier with one broad rule.
For each rule, review an expected match, an ordinary non-match, and an ambiguous example. Record why the detector is useful and who should investigate its findings. Regular expressions recognize patterns, so the surrounding business context still matters.
Keep reviews tied to the page version
Mark as reviewed records an attestation only when the caller has update permission, the page has a completed scan, and the saved scan version matches the current Confluence version. A later edit means that the previous review no longer represents the current version. The current build sets the next-review date to 90 days after attestation.
This version link prevents an old approval from being treated as a review of newly changed content. It does not certify compliance or guarantee that all sensitive data was detected. Reviewers should examine coverage warnings before recording their decision.
For teams that also need editors to explain changes, Ravenshift’s Page Change Comments for Confluence addresses edit reasons separately. A change reason explains why content changed; a scanner review records assessment of a particular version. Keep both records attached to the decision they actually support.
Dashboard access and evidence exports
Confluence administrators can use the global dashboard and approve specific users as dashboard viewers. Backend authorization enforces this access. Page-scoped results and audit events are filtered through the viewer’s current Confluence visibility; dashboard approval does not grant access to otherwise inaccessible pages.
Scan and configuration data is stored in tenant-scoped Forge SQL. CSV and JSON exports are bounded and permission-filtered. CSV output neutralizes spreadsheet-formula prefixes. Treat exports as app-generated evidence, and review the exported scope before sharing it with another team.
A useful evidence record states the assessed scope, scan version, review status, and incomplete coverage. Preserve the distinction between a current result and an older export: an exported snapshot does not update when a page is edited later.
What the scanner does not assess
The current build does not classify anonymous, guest, public-link, restricted, or space-wide exposure. Exposure is reported as not assessed. Use Confluence’s access-management tools for permissions review. Ownership context may be available, but the app does not classify owners as active, deactivated, internal, or external.
Lifecycle signals are advisory. The app does not automatically archive pages, change restrictions, edit page bodies, or add lifecycle labels. Its Archived governance classification is not a native Confluence archive operation.
Use the scanner for pattern detection and documented review within its supported scope. Combine that evidence with content ownership, access review, and the remediation process your organization uses. Before deployment, confirm the released version and privacy disclosures rather than assuming every future feature is already available.
Plan the release and ongoing review
PII & Secret Scanner for Confluence is coming soon from Ravenshift. Contact [email protected] for release availability. No public Marketplace installation link is claimed on this page yet.
Choose a pilot space, agree on finding ownership, and review incomplete coverage before expanding the rollout. Revisit detector settings when documentation formats change, and check the current product documentation for release-specific limits. The aim is a repeatable review process with clear evidence, not a dashboard count detached from the documents behind it.