Learn how Scroll Sites optimizes your site for the best possible ranking on search engines.
If, when and how your site is included in the search results of web search engines depends on a series of factors. Some of these factors are controlled by the search engines but some can be controlled by you as a user.
Is my Scroll Site Optimized For Search Engines (SEO)?
We don’t prevent search engines from finding, crawling and indexing your public Scroll site, unless you choose to protect your site with a login.
That means each page of your publicly-available site will be available for search engines to crawl. There is no option to prevent search engine indexing with Scroll Sites.
Once crawled and indexed, SEO-wise Scroll sites perform better than public Confluence pages. As static HTML sites, Scroll sites are more focused on content than public Confluence pages, which load a lot of content-unrelated page elements.
We also take great care to deliver your site in the most search engine optimized way possible (e.g. by generating a sitemap and meta descriptions for pages).
Using a Sitemap and Robots.txt To Help Search Engines Crawl Your Site
With each Scroll site, we automatically generate a sitemap.xml and robots.txt file with the names /sitemap.xml and /robots.txt in your domain's top level directory (e.g. example.scroll.site/).
A sitemap is a XML file that lists the URLs for a site. Scroll Sites refreshes your sitemap when you update your site, so newly published pages appear after your next site update.
A robots.txt tells search engines which URLs they should access and crawl on the site. In combination, they allow search engines to crawl the site more intelligently.
The robots.txt file generated by the app instructs search engines to access only those pages that are listed in your sitemap. Search engines will automatically look for these files in their location.
Please note that you cannot modify the robots.txt file currently.
Canonical Link Elements for All Articles and Global Pages
Scroll Sites automatically reports canonical URLs for all pages of your help center. This effectively improves search engine optimization (SEO) for your Scroll site, as duplicate content that is available from multiple URLs stops being an issue for search engines.
For example, you won't see /contentsource and /contentsource/index.html flagged as duplicates because /contentsource/index.html is now marked as the original source of the content.
Factors Affecting Site Indexing
We can’t determine how long it takes for different search engines to index your site and all its content. How long indexing takes can be impacted by:
-
The specific search engine (and their crawling and indexing approach)
-
The theme and content choices you make for your Scroll site
-
The proactive SEO steps you decide to take in order to be indexed by a search engine
Use Google Search Console To Appear in Google Search Results
You can take a series of proactive steps to control a search engines' ability to find and parse your content.
To be included in the Google index, you can use Google Search Console to submit an indexing request for your Scroll site. Google has extensive help resources on this topic.
Troubleshooting:
1. Google Search Console Shows No Pages Indexed
If Google Search Console reports that none of your pages are indexed, or your pages stay at "Discovered – currently not indexed," work through the following checks in order.
Give it time
Indexing isn't instant. After you launch a new site or publish a major update, Google can take days to weeks to crawl and index your pages. Small sites with few inbound links take longer. If your site is new, wait 2 to 4 weeks before you investigate further.
Confirm your site is publicly accessible
Google can only crawl pages that are publicly available, so start here:
-
Check that no token or login protects your site. Both token protection and SAML single sign-on make your site uncrawlable, so a site that needs to rank in search results has to stay open to everyone. A site that requires authentication never appears in search results.
-
Open your site in a private browser window and confirm you can reach your pages without signing in.
Check your robots.txt
Open https://your-site-domain/robots.txt in a browser. Confirm that it references your sitemap, and that no unintended Disallow directive blocks the pages you expect Google to index.
Check your sitemap.xml
Open https://your-site-domain/sitemap.xml and confirm that it lists the pages you want indexed. A sitemap that carries many archived or older versions can dilute Google's crawl priority. See "Reduce the number of URLs you advertise" below.
If your browser says the XML file has no style information
When you open /sitemap.xml or a sub-sitemap in a browser, you may see this message above the content:
This XML file does not appear to have any style information associated with it. The document tree is shown below.
This is normal and it isn't an error. Sitemaps are raw XML files written for search engines, not for people, so they carry no styling. Your browser is telling you it's displaying the raw document tree instead. If you can see your page URLs listed below the message, your sitemap is working.
2. Google Search Console Shows “Could Not Fetch” Status
If the Sitemaps report in Google Search Console shows a "Could not fetch" status, check whether your site includes a content source that has no indexable pages.
Scroll Sites builds your sitemap from every content source in your site. When a content source contains no indexable pages, Scroll Sites still generates a section for it, but that section lists no URLs. Google can report "Could not fetch" for a sitemap that contains an empty section.
A content source ends up with no indexable pages when every page in it carries an exclusion label, such as scroll-sites-only-url.
If you want a page gone from your site entirely, use the scroll-sites-no-publish label instead. The page URL then returns a 404, so any inbound links to it stop working. Choose scroll-sites-only-url when you want to preserve those inbound links, and scroll-sites-no-publish when the content should not be reachable at all.
Neither label leaves the page available to AI search. AI search answers from indexed page content, and neither label allows the page content to be indexed.
To find the content source that causes this:
-
Open
https://your-site-domain/sitemap.xmlin a browser. -
Look for a section that contains no URLs.
-
Match that section to a content source in your site.
To fix it, choose one of these options:
-
Remove the empty content source from your site. See Manage Content for a Site.
-
Remove the exclusion label from at least one page in that content source, so the source contributes at least one URL. See Exclude Articles From Navigation and Search.
Then update your site and resubmit your sitemap.
Submit and check your sitemap in Google Search Console
Follow these steps in Google Search Console:
-
Open Google Search Console for your site's domain.
-
Go to Sitemaps and submit your sitemap URL if you haven't already.
-
Open the Pages report and review the coverage status.
Two statuses in that report point to different problems:
|
Status |
What it means |
|---|---|
|
Discovered – currently not indexed |
Google found the URL but hasn't visited it yet. This often points to a crawl budget problem, where too many URLs compete for limited crawl resources. |
|
Crawled – currently not indexed |
Google visited the page and chose not to include it. This usually relates to content quality signals, or to a high proportion of similar content across the domain. |
Rule out label and permission exclusions
Scroll Sites excludes pages that carry the scroll-sites-only-url label from your sitemap and from llms.txt. If that label sits on pages you want indexed, remove it and update your site.
Also check that no Confluence restriction stops a page from reaching your live site. Both space permissions and page restrictions can prevent publication.
Reduce the number of URLs you advertise
Sites with many archived versions face a specific problem. The volume of near-duplicate content can lead Google to deprioritize the whole domain, even when canonical tags are correct.
To narrow the set of URLs you advertise to search engines:
-
Add the
scroll-sites-only-urllabel to the root page of each archived version you want to keep online but exclude from search results. The label cascades to every child page. -
Update your site, then confirm those URLs no longer appear in
/sitemap.xml. -
Resubmit your sitemap in Google Search Console.
-
Use URL Inspection > Request Indexing for your most important current pages.
Pages you exclude with this label stay reachable through their direct URL, so inbound links keep working. However, the label also removes them from your site navigation and internal search. Excluding pages from the sitemap while keeping them visible in navigation isn't supported yet.
After you make changes
Once you've corrected any access, exclusion, or sitemap problem:
-
Update your site.
-
Resubmit your sitemap in Google Search Console under Sitemaps > Resubmit.
-
Use URL Inspection > Request Indexing on your key pages to prompt a re-crawl.
-
Allow 1 to 2 weeks for the changes to take effect.
Related Pages