Key takeaways
- A desktop crawler already reaches network-restricted staging (VPN, localhost, IP allowlist) because it runs from your machine; a cloud crawler has to be allowlisted first
- Since Crawl Cove 1.2.0 a password-gated staging site works too: give the crawl a username and password for an HTTP Basic Auth prompt, or a login page URL plus two CSS selectors for a form, and it stays signed in for the whole pass
- Credentials are encrypted per site with remember, use-once and forget options, and are only ever sent to the staging site's own origin, never to a redirect target
- Most staging hosts ship a blanket Disallow in robots.txt; tick Ignore robots.txt for that crawl, on sites you own, or the crawl stops at page one
- Third-party single sign-on, magic links and two-factor prompts are not a login form the crawler can fill, so those still need network access or a test account without them
Yes. A staging site behind a login can be crawled, and since Crawl Cove 1.2.0 that includes staging sites gated by a password rather than only by the network. This post covers which kinds of gate a desktop crawler can get through, how to set up a crawl that signs in, and the handful of things that are only ever wrong on staging and are worth checking before the site goes live.
Which kind of gate is yours
"Behind a login" means three different things, and the answer differs for each.
- Network-restricted. The site is on a VPN, on localhost, or behind an IP allowlist, and once your laptop is inside there is no password at all. A desktop crawler has always handled this, because the crawl runs from your machine and reaches whatever your browser reaches. A cloud-based site audit has to be allowlisted by IP or user agent first, which is why the comparison pages on this site keep returning to it (Screaming Frog vs Semrush is the clearest case).
- HTTP Basic Auth. The browser pops up its own username and password box before the page loads. This is the most common staging gate on agency hosting, WP Engine, Kinsta and similar, and since 1.2.0 Crawl Cove signs in for you.
- A login page. The site shows its own form, sets a session cookie when you submit it, and every page after that is checked against the cookie. Also supported since 1.2.0, with slightly more setup because the crawler has to be told where the form is.
The one that still does not work is a gate the crawler cannot complete: single sign-on that redirects to Google, Microsoft or Okta, a magic link sent by email, or a two-factor code. A form login fills a username and password field on a page that belongs to the staging site. It cannot click through a third-party identity provider, and it should not try, because the credentials must never leave the staging site's own origin. For those, use a test account without the extra step, or fall back to network access.
Crawling behind HTTP Basic Auth
In Crawl Cove, start a new crawl as normal, then open the Staging login section of the form. It asks for two things: the username and the password. That is the whole setup.
The credentials are sent as an Authorization header on every request whose origin matches the URL you started the crawl from, and on nothing else. If the staging host redirects a request off-site, to a CDN, a tracker or an external login provider, that redirect target does not receive the header. This scoping is deliberate and is the reason the login is tied to a site rather than to the crawl in general.
Three checkboxes control what happens to the password afterwards:
- Remember for future crawls of this site keeps the pair, encrypted with the operating system's own credential store, so the next crawl of that site needs nothing typed.
- Leaving it unticked uses the credentials once and discards them. Nothing is written to disk that you did not agree to.
- Forget the saved login for this site appears once a login is saved, and removes it.
The onboarding wizard's first crawl accepts a staging login too, so a staging site can be the first thing a new install audits.
Crawling behind a login page
Open Staging login, form-based instead. This one needs the login page URL and two CSS selectors: one for the username or email field and one for the password field. A submit selector is optional; without it the crawler presses Enter in the password field, which is what most forms expect.
For a WordPress staging site the login page is /wp-login.php, the username field is #user_login, the password field is #user_pass, and the submit button is #wp-submit. For anything else, right-click the field in your browser, choose Inspect, and copy the id or name attribute: #email or input[name="password"] are typical.
Two things are different from the Basic Auth case and worth knowing before you rely on it:
- It needs a browser. The login is performed by a real headless browser, once, at the start of the crawl, and the session cookie it receives is then reused for the whole pass. The machine needs Google Chrome or Microsoft Edge installed, or the optional Playwright Chromium. An HTTP Basic Auth login needs none of that.
- A failed login does not stop the crawl. If the selectors are wrong or the site rejects the credentials, the crawl continues unauthenticated, exactly as if no login had been configured, and the failure is logged rather than thrown. So check the result: a staging crawl that returns one page, or a run of pages that are all the login form, means the login did not take. Fix the selectors and run it again.
The cookie is scoped the same way the Basic Auth header is: it is sent to the staging site's own origin and to nothing else. The same remember, use-once and forget controls apply.
The robots.txt trap on staging
Most staging crawls that "do not work" fail for a reason that has nothing to do with the login. Staging hosts ship a robots.txt with a blanket Disallow: /, often one you cannot edit because the host injects it, and Crawl Cove honours robots.txt by default. The crawl fetches the file, reads the rule, and stops.
The New crawl form has an Ignore robots.txt checkbox for exactly this case. It is off by default and applies to that crawl only. The raw robots.txt is still fetched and reported, so the AI-crawler and robots checks still run; only the gating is bypassed. Use it on sites you own or are authorised to audit, which describes every legitimate staging crawl.
The same rule applies to the open-source CLI: its --ignore-robots flag exists for staging hosts, as the terminal walkthrough explains. The CLI has no credential option, though, so it covers network-restricted staging and not a password gate.
What to check on a staging crawl
A staging crawl is not a rehearsal of the production crawl. It has its own findings, and most of them are things that are correct on staging and become wrong the moment the site is copied to production. Look for these before sign-off:
- noindex on every page. Expected on staging, catastrophic if it ships. Crawl Cove reports it as a finding either way; on staging, what you want is the list, so you can confirm the live site is clean after launch.
- Canonical tags pointing at production. A staging copy of a live site often carries canonicals to the live URLs. Fine for a preview, but it means any page you add on staging has a canonical that points at a URL that does not exist yet.
- Absolute links to the production host. If the crawl finds only a handful of pages on a large site, the internal links probably point at the live domain and the crawler, correctly, treated them as external.
- Hreflang and sitemap URLs on the wrong host. Both are usually generated with the site URL and both silently break when that URL is the staging one.
- Broken links to assets that only exist in production. Uploads directories are the usual culprit.
Then crawl the live site after launch and compare. Crawl Cove keeps each crawl as a versioned run and works out what is new, fixed, still open and regressed between two of them; how to compare two SEO crawls walks through reading that diff.
How other crawlers handle it
Screaming Frog has supported this for years: its user guide lists basic and digest authentication, detected automatically when a crawled page asks for it, and a forms-based mode for login pages. If you are moving between the two, the mechanism is the same shape, and a Frog crawl of the staging site can be converted into the Crawl Cove export format if you want the two side by side.
Cloud site audits are the ones to plan around. Because the crawl runs from the vendor's servers, a staging site has to be reachable from the internet and the vendor's bot allowlisted, and a password gate needs the vendor to support it explicitly. That is the practical reason a desktop crawler is still the tool most teams use for pre-launch checks, and the SEO crawler page covers the rest of what a local crawl gets you.
Frequently asked questions
- Can Crawl Cove crawl a staging site behind a login?
- Yes, since version 1.2.0, released on 27 September 2026. In the New crawl form, open Staging login and enter a username and password for an HTTP Basic Auth prompt, or open the form-based section and give it the login page URL and the CSS selectors of the username and password fields. The crawler signs in once and stays signed in for the whole crawl.
- Does a staging login work with single sign-on or two-factor authentication?
- No. Form login fills a username and password field on a page that belongs to the staging site and submits it. A gate that redirects to Google, Microsoft or Okta, sends a magic link, or asks for a one-time code is not something the crawler can complete. Use a test account without those, or reach the staging site over a VPN or IP allowlist instead.
- Where are the staging credentials stored?
- Encrypted on your own machine, per site, using the operating system's credential encryption. Tick Remember for future crawls of this site to keep them, leave it unticked to use them once, and tick Forget the saved login to remove them. They are sent only to requests on the staging site's own origin, never to an external redirect target such as a CDN or a login provider.
- Why did my staging crawl stop after one page?
- Almost always robots.txt. Staging hosts commonly ship a blanket Disallow for every crawler, and Crawl Cove honours it by default. Tick Ignore robots.txt in the New crawl form for that crawl only, which is appropriate for a site you own or are authorised to audit.
- Do I need anything extra for a form-based login?
- A browser. The form is filled by a real headless browser, so the machine needs Google Chrome or Microsoft Edge installed, or the optional Playwright Chromium. An HTTP Basic Auth login needs nothing extra.
- Can the Crawl Cove CLI crawl behind a login?
- Not yet. The open-source crawlcove-cli has no credential option; it has an ignore-robots flag for staging hosts you own, which covers network-restricted staging but not a password gate. Use the desktop app for a gated site.