What the web check
sends, and what it cannot see.
This page has two readers. One is looking at a report and wants to know what it means. The other found this address in their access log, because this is what the request said it was, and wants to know exactly what reached their server. The second question is answered first.
What was sent to your server
If you are here from a log line, this is the whole of it. There is nothing else to find, and none of it is a probe for a weakness.
Every request names itself. The
User-Agent header reads
porch/1 (+https://porch.denyfirst.dev/web/method), from every installation of the tool, so a
line in your log leads here.
One GET of /, over HTTPS and over
plaintext, and one of /.well-known/security.txt.
Nothing else was requested, and no address is constructed by this program
at any point: there is no attempt at /admin, no
/backup.zip, and no second guess of any kind. This reads
what a server volunteers to every visitor.
And one GET of / on the other form of
the name. A name without www. gets it added, a name
with it gets it removed, and that one name is asked for its root over
HTTPS. What it answers is not followed: a redirect says where it
is sending a visitor in its own header, so following it would buy another
connection to a server and no new fact.
Nothing else is looked up, and no name is searched for.
Four characters are added or removed from the name you gave. There is no
attempt at mail, dev, staging or
old — a tool that went looking for those would be
enumerating an estate, which is a different instrument with a different
argument for existing. The report names the form it compared, so you can
see which question was answered. It does not claim to have found your
apex: telling example.co.uk from blog.example.com
needs the public suffix list, a third-party list this project does not
carry and will not write its own copy of.
Why it is worth a connection. This is the question you
cannot ask of your own site, because your habit answers it for you:
whoever types the short form every day never finds out what a visitor
typing www gets, and whoever bookmarked www
never finds out what happens at the short form. Both are in every
visitor's muscle memory and usually only one has ever been tried. The
commonest thing it finds is two names serving two sites with nothing
joining them — one of which stopped being updated some years ago. Nothing
here is graded: no document requires a www form to exist, or
to redirect anywhere.
And one handshake, carrying no request, to an address the name
publishes for IPv6. It is opened, the certificate is read, and
it is closed; nothing is asked of the server. This exists because a scan
could not answer the question otherwise. Reaching a site means trying
its addresses until one answers, which is the right way to reach it and
the wrong way to measure it: a host with a working IPv4 address and an
AAAA record pointing at nothing answers every scan, reads
perfectly, and cannot be opened at all from a network that has only the
newer protocol. The operator is the last to find out, because their own
machine has both. Where the name publishes no AAAA record,
no connection is made and the report says so — that is a decision
somebody made and not a fault. Nothing here is graded either: no
document requires a site to be reachable over IPv6.
These three run last, and say so when the scan runs out of time. The security contact, the IPv6 address and the other form of the name are measured after the pages, sharing one time budget with them. A slow site spends it, and then these are what go without — so each of them says the scan ran out of time before this was measured rather than reporting what it saw. A file that could not be fetched because this scanner gave up is not a site with no security contact, and printing the first as the second would be this program describing its own clock as a fault in your server. Once the time is gone nothing further is looked up or connected to, either: a request that cannot finish is a line in your log bought for nothing.
A guess and a registered address are different things.
This paragraph read no guessing under
/.well-known until September 2026, and that was wrong
in a way worth correcting in public rather than quietly: the space is a
register, not a hiding place. RFC 8615 created it for addresses a server
offers to anyone who asks, and every entry in it is defined by its own
document. RFC 9116 defines the one this check fetches:
security.txt exists so that a stranger who finds a fault in
your site knows who to tell, which makes reading it the clearest case
there is. Two other checks fetch one each — a mail policy, and only where
the zone published a record saying it is there, and the file half of
proof of control, which whoever publishes it is asking to have read.
Three addresses, named in one list in the source, each beside the
document that defines it, so that a fourth cannot arrive without anybody
noticing.
What is kept from security.txt is a count and a
date. The report says how many ways to report a fault the file
names and when the file says it expires — never the addresses
themselves. The state worth seeing is the one that is neither present nor
absent: a file published three years ago and expired since, which tells
whoever found a problem that they are expected somewhere nobody is
waiting. Nothing about this is graded. No document requires the file of
anybody, and a verdict here would be a threshold invented by this
project.
Redirects were followed, up to five, and only where a
Location header named. Each address after the
first was chosen by your server rather than by this scanner. A
Location in another scheme is not followed, an overlong one
is not followed, and credentials in one are stripped before the request
is made. Where the chain stopped, the report says whether it ended or
was declined, because a reader who cannot tell those apart cannot read
the chain at all.
A redirect off what this deployment may reach is not
followed. Whatever decides which hosts an installation of this
tool connects to is asked again about the address a
Location header names, before anything is dialled — because
that address was chosen by a server rather than by the person running
the scan.
On this deployment that means the hosts this project owns,
so a redirect away from them ends the chain and the report says
so.
Whether the body was read depends on which installation reached you, so this says both. The user agent names this address from every installation of the tool, so one flat sentence here would be true of one of them and false of the one in your log.
This deployment reads the page, and only its own. It connects to nothing but the hosts this project owns, so a page it reads is ours: once, to a bound of one megabyte, streamed, and only the page at the address whose headers it already fetched. A request in your log from this deployment would mean your server is one of ours; any other installation of the tool is somebody else's, and the next paragraph is about those.
An installation somebody runs themselves may read the page — once, to a bound of one megabyte, streamed, and only the page at the address whose headers it already fetched.
Where a page is read, nothing of it is kept. No link
is followed, no script retrieved, and nothing else on your server is
asked for. What survives the read is a short list of facts — whether a
Content-Security-Policy was declared in the markup, and the
host names of what the page pulls in from somewhere other than itself:
anything loaded over plaintext, and anything loaded from another origin.
A page that loads only its own things from its own host produces an
empty list, because nothing asks about that.
A host, and never an address. The path is dropped, the
query is dropped, and any credentials in it are dropped before anything
else — http://user:token@host/ in somebody's markup is a
credential, and a report carrying it would publish it to everyone the
report is shown to. There is no field anywhere in a report that can hold
markup, in the same way there is no field that can hold a cookie's
value. The reason for reading the page at all is that mixed content, a
script from elsewhere with nothing pinning it, and a policy declared
with <meta http-equiv> are ordinary things a site has
and nobody can see from headers; the reason nothing else is kept is that
a page holds keys, tokens and names, and a report is a thing people
paste into issue trackers.
Only the headers this check grades were kept. An allow list, not a deny list. Whatever else your server sent — internal host names, software versions, request identifiers — was discarded where the response was parsed rather than filtered out of the report later.
No cookie value was recorded, and there is nowhere to put one. The report holds a cookie's name and the attributes that decide whether it is safe. There is no field for a value at all: a value is a session identifier as often as not, and a report is a thing people paste into issue trackers.
Ports 80 and 443 only, and nothing was sent to any address that is private, loopback, link-local or reserved — including an address reached by following one of your own redirects.
If you would rather this domain were not scanned at all, that is arranged by asking and needs no explanation. The address, and what else a person here will answer, are under stopping a scan.
How to read a report
A report says four kinds of thing, and telling them apart is most of knowing what it means.
What the site sends is listed whether or not it declares it. Thirteen response headers are read on every scan and three of them are graded; the rest — the content policy, the referrer policy, the framing and cross-origin policies — are facts about what a browser is told, and no document requires any of them. Each is shown with what it said, or with none where the response carried nothing: a row that is not there reads as a question nobody asked. Beside them: where a content policy was declared, how many cookies were set and with which attributes, and what the page pulls in.
Findings are graded. Each one names a rule, states what was measured, and cites the document it rests on. Grading is worst-case: a site reached in the clear is reached in the clear however sound its policy declaration is.
Observed is what the check established and deliberately
does not grade. The length of an HSTS max-age
is the clearest case: no standards body publishes a minimum, and OWASP
explicitly recommends a short one during a rollout, so the value is
described in years or days along with what a browser does when it lapses
— and never failed against a threshold this project invented. The same
holds for a missing includeSubDomains on a host with
nothing beneath it, and for a temporary redirect from the plaintext
address, which works.
Not established for this host is what this particular scan could not settle, for a reason that lies with the server in front of it. These qualify that report's verdict and travel with it.
Limits of this method are below. They are true of every scan this check runs, so they are stated once, here, rather than repeated on every report where they read as faults of the server being looked at.
Two distinctions decide more than they look. A host that answered
nothing is ungraded, not weak: declaring no policy is
a claim about a server, and nothing was ever spoken to. And a host with
nothing listening on port 80 is the safest arrangement there is, which
is not the same fact as a host answering 200 in the clear,
though both produce an empty list of headers.
The policy graded is the last one carried by a hop made over TLS, because that is the one a browser would end up holding. Reading the last hop describes a policy no browser holds the moment a site downgrades at the end of its chain; reading the first describes one a later hop replaced.
Only the root was asked
One request was made, to the root of the site, and no other address on it was asked for. Another page may answer with different headers and load different things, and nothing here describes any page but this one.
No browser ran here
Nothing was executed. Where the page was read, what it says it loads was read out of the markup — so anything a script fetches once it runs, and anything assembled after the page arrives, was not seen. Whether a declared policy is enforced in practice is visible only to a browser, and no browser ran here.
One answer, from one machine, at one moment
A name served by several machines, or one running an experiment, can answer the next visitor differently. This is what one address said once.
What else is written down
This check has its own rule set, separate from the one that grades a TLS handshake, because they answer different questions over different evidence and a single name for both would make two incomparable things look comparable. What the other one can and cannot see is on its own page.
Every property this project holds itself to is in docs/invariants.md, each one naming the test that guards it and, where there was one, the mistake that produced it. What changed between rule sets is in docs/policy-changes.md, because a report graded under one version is not comparable with a report graded under another.
What is recorded about a scan here, and how to have a domain excluded, are on the privacy page.