{"id":1338,"date":"2026-09-21T08:00:00","date_gmt":"2026-09-21T13:00:00","guid":{"rendered":"https:\/\/www.seomos.com\/?p=1338"},"modified":"2026-09-10T02:42:26","modified_gmt":"2026-09-10T07:42:26","slug":"robots-txt","status":"publish","type":"post","link":"https:\/\/www.seomos.com\/en\/blog\/robots-txt\/","title":{"rendered":"Robots.txt: How to configure it without blocking important pages"},"content":{"rendered":"<div class=\"seomos-resumen\">\n<p><strong>Robots.txt is a text file located in the root of a host that tells compatible crawlers which routes they can request.<\/strong> Its primary function is to manage crawling, not to remove pages from the index or protect information. A blocked URL may appear in results if other sites link to it; for private content, it corresponds to authentication, and to exclude a crawlable page, it may correspond to... <code>noindex<\/code>.<\/p>\n<\/div>\n<p>A single incorrectly implemented line can block a section or the entire site from search engines. That&#039;s why robots.txt should be treated like a production configuration: with a target, owner, tests, history, and recovery plan.<\/p>\n<h2 id=\"que-es\">What is the robots.txt file?<\/h2>\n<p>It implements the Robots Exclusion Protocol. It groups rules for one or more agents and path patterns. Recognized search engines usually respect them, but a malicious actor can ignore them. The file is public, and anyone can read the specified paths.<\/p>\n<p>If there is no rule blocking a URL, access is considered allowed. Having an empty file or a rule <code>Allow: \/<\/code> It does not improve positioning; it only makes access explicit.<\/p>\n<h2 id=\"ubicacion\">Where it should be and what applies<\/h2>\n<p>It is served as <code>https:\/\/ejemplo.com\/robots.txt<\/code>. It doesn&#039;t work in <code>\/folder\/robots.txt<\/code>. Its scope is limited to protocol, host, and port: the file of <code>www<\/code> It does not control a subdomain, and HTTPS does not inherit rules from HTTP.<\/p>\n<p>Each relevant host needs its own decision. Check the canonical domain, subdomains, staging, and CDN. The file must respond publicly and in UTF-8 text; HTTP redirects and errors can alter how it&#039;s interpreted.<\/p>\n<h2 id=\"sintaxis\">Basic syntax<\/h2>\n<p><code>User-agent<\/code> Start a group. <code>Disallow<\/code> restricts a route; <code>Allow<\/code> can open an exception within a locked pattern. <code>Sitemap<\/code> declares an absolute URL for the sitemap. The character <code>#<\/code> Start a comment.<\/p>\n<p>Routes are case-sensitive. Google supports wildcards as defined in its documentation, but other crawlers may interpret them differently. Design for the standard first and validate with the preferred crawlers.<\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" src=\"https:\/\/www.seomos.com\/wp-content\/uploads\/2026\/09\/robots-txt-model.webp\" alt=\"Rutas f\u00edsicas representan reglas de acceso y ubicaci\u00f3n de sitemap para rastreadores\" width=\"1200\" height=\"800\" decoding=\"async\"><figcaption>The rule should be evaluated against the agent and the complete route that is actually being requested.<\/figcaption><\/figure>\n<h2 id=\"ejemplo\">Minimal example with commentary<\/h2>\n<pre><code>User-agent: * Disallow: \/internal-search\/ Allow: \/ Sitemap: https:\/\/www.example.com\/sitemap.xml<\/code><\/pre>\n<p>The example allows the site but blocks a specific path. It should not be copied without reviewing the site architecture. If the internal search engine returns useful pages or the path changes, the rule would have a different effect.<\/p>\n<p><code>Allow: \/<\/code> It is redundant when there is no broader lock-up. Clarity matters: a few justified rules are safer than an ownerless historical collection.<\/p>\n<h2 id=\"rastreo-indexacion\">Crawling is not indexing<\/h2>\n<p>Crawling means requesting a resource. Indexing means processing it so it can appear in search results. Robots.txt controls the first step. Google warns that a blocked URL can still be indexed without a description if discovered through links.<\/p>\n<p>If you need to remove a public results page, allow tracking and delivery. <code>noindex<\/code> until it is processed, or remove it with an appropriate status. If it contains private information, require authentication; do not rely on voluntary instruction.<\/p>\n<h2 id=\"canonical\">Robots.txt does not canonicalize<\/h2>\n<p>A canonical tag requires the search engine to be able to read the HTML or header. Blocking the variant can prevent that signal from being processed. For duplicates, use <a href=\"https:\/\/www.seomos.com\/en\/blog\/canonical\/\">canonical tag<\/a>, redirects, consistent links and sitemap as appropriate.<\/p>\n<p>Robots can reduce access to low-value spaces, but they don&#039;t replace architecture. First, they prevent the creation of infinite combinations, and then they control the inevitable.<\/p>\n<h2 id=\"seguridad\">Robots.txt is not security<\/h2>\n<p>The file publicly advertises paths and does not enforce them on all clients. Panels, documents, and private environments require authentication, authorization, and server configuration. Removing a path from the file does not make it secret.<\/p>\n<p>Also check backups, storage, and shared URLs. If there was exposure, correct access and assess the incident; editing robots.txt does not revoke existing copies or links.<\/p>\n<h2 id=\"recursos\">CSS, JavaScript, and images<\/h2>\n<p>Blocking necessary resources can prevent a search engine from rendering and understanding your page. Don&#039;t block entire folders of themes or scripts without checking what the templates consume. Google recommends allowing resources relevant to rendering.<\/p>\n<p>For images or video, the rules can influence file crawling and resource visibility, but the containing page is a separate object. Separate the objectives and test both.<\/p>\n<h2 id=\"ecommerce\">Parameters, filters and ecommerce<\/h2>\n<p>Ecommerce platforms generate order, filters, search capabilities, and tracking. Each parameter is categorized by content changes, demand, and links. Robots can restrict access to vast areas, but if they are already indexed, blocking them can freeze outdated signals.<\/p>\n<p>Reduce URL generation, normalize URLs, and link selected facets. Align canonical tags, sitemap, and navigation. Don&#039;t use a global rule for all URLs. <code>?<\/code> without checking pagination and legitimate functions.<\/p>\n<h2 id=\"wordpress\">Robots.txt in WordPress<\/h2>\n<p>WordPress can generate a virtual file, and SEO plugins can modify it. Check the public search results, not just the editor. The option to discourage search engines may result in broad blocking or meta robots blocking, depending on the version and configuration.<\/p>\n<p>In production, verify after migrating, cloning, or changing the domain. Staging environments must be protected with access controls, not just robots.txt controls. Never replace server files without knowing if a virtualized layer or CDN is in place.<\/p>\n<h2 id=\"bloqueos-utiles\">When can a lock be useful?<\/h2>\n<ul>\n<li>Internal searches with no indexable value.<\/li>\n<li>Technical combinations that create enormous spaces.<\/li>\n<li>Endpoints or resources that are dispensable for compliant tracing.<\/li>\n<li>Navigation parameters without distinctive content, evaluating previous state.<\/li>\n<li>Administration areas, as a complement and not as security.<\/li>\n<\/ul>\n<p>The benefit is weighed against the risk of hiding resources, links, or signals. On small sites, a complex bot rarely pays off.<\/p>\n<h2 id=\"errores-sintaxis\">Syntax and scope errors<\/h2>\n<p>A <code>Disallow: \/<\/code> Blocks everything for the group. A missing slash can change scope. Rules are case-sensitive. Repeated groups, spaces, and unsupported directives lead to incorrect assumptions.<\/p>\n<p>Another mistake is editing the wrong host file or trying an administrative session that uses a different path. Always retrieve the exact public URL and check the status, content, and headers.<\/p>\n<h2 id=\"auditoria\">How to audit robots.txt<\/h2>\n<ol>\n<li>Download files from each relevant host and protocol.<\/li>\n<li>Logs HTTP status, redirection, and content.<\/li>\n<li>List agents and groups.<\/li>\n<li>Test strategic URLs and patterns.<\/li>\n<li>Compare with crawling, sitemap, and links.<\/li>\n<li>Review rendering resources.<\/li>\n<li>Relate rules to a documented objective.<\/li>\n<li>Remove or correct by controlled change.<\/li>\n<\/ol>\n<p>A <a href=\"https:\/\/www.seomos.com\/en\/seo\/bogota-seo-audit\/\">SEO audit<\/a> You should test real pages: homepage, categories, products, articles, pagination, languages, resources, and responsive paths.<\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" src=\"https:\/\/www.seomos.com\/wp-content\/uploads\/2026\/09\/robots-txt-process.webp\" alt=\"Cuatro carriles separan rastreo indexaci\u00f3n consolidaci\u00f3n y control de acceso\" width=\"1200\" height=\"800\" decoding=\"async\"><figcaption>Crawling, indexing, canonicalization, and security require different controls.<\/figcaption><\/figure>\n<h2 id=\"cambio\">How to deploy a secure change<\/h2>\n<p>Save a copy, explain the objective, and prepare pre- and post-tests. Evaluate first in an environment that won&#039;t affect production, but remember that the staging host has its own scope. Schedule log monitoring and Search Console tracking.<\/p>\n<p>After posting, request <code>\/robots.txt<\/code> Without caching, test URLs and confirm that the CDN and server are delivering the latest version. Google may keep a temporary copy; use official mechanisms where appropriate and monitor.<\/p>\n<h2 id=\"incidente\">What to do in case of an accidental blockage<\/h2>\n<ol>\n<li>Confirm the public file and the affected agent.<\/li>\n<li>Identify change, scope, and time.<\/li>\n<li>Restore a known version or correct the rule.<\/li>\n<li>Purge only the necessary cache.<\/li>\n<li>Verify from outside and test URLs.<\/li>\n<li>Review robot and inspection report.<\/li>\n<li>Monitors crawling and indexing.<\/li>\n<li>Document cause and prevention.<\/li>\n<\/ol>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" src=\"https:\/\/www.seomos.com\/wp-content\/uploads\/2026\/09\/robots-txt-example.webp\" alt=\"Especialista diagnostica un bloqueo masivo de p\u00e1ginas provocado por robots.txt\" width=\"1200\" height=\"800\" decoding=\"async\"><figcaption>Restoring the file is the first step; the recovery should be verified in crawling and indexing.<\/figcaption><\/figure>\n<h2 id=\"medicion\">What to measure<\/h2>\n<p>Observe requests by agent and section, errors, important blocked pages, uncrawled resources, and indexing of known URLs. Traffic changes are not automatically attributed to robots; relate date, pattern, and evidence.<\/p>\n<p>On large sites, it analyzes logs. On all sites, it retains canonical evidence. The goal is not to &quot;block more,&quot; but to facilitate the discovery of useful content by permitted crawling without overloading systems.<\/p>\n<h2 id=\"codigos-http\">What happens when the file fails?<\/h2>\n<p>Behavior can vary depending on the HTTP status and the time of the failure. A successful response should deliver the expected file; a 404 error is usually interpreted as a lack of restrictions, while server errors may cause temporary caution. Consult the current crawler documentation for each case.<\/p>\n<p>Don&#039;t design SEO availability around an assumption. Monitor the endpoint, maintain a secure minimum version, and prevent robots from relying on an unstable application. A cross-host redirect can also leave rules unenforced even if the browser displays content.<\/p>\n<h2 id=\"agentes\">Specific agents and groups<\/h2>\n<p>A group for <code>*<\/code> It covers compatible agents not specifically addressed. If you create particular groups, check which rules actually combine or select. Don&#039;t assume that typing a company name covers all of its trackers.<\/p>\n<p>Document why an agent is being treated differently, who requested the change, and how it will be evaluated. Blocking tool trackers can affect audits; allowing user-activated clients can have other implications. Names and behaviors evolve, so check official references.<\/p>\n<h2 id=\"logs\">Validation with server logs<\/h2>\n<p>The logs show actual requests, declared agent, route, response, and time. Use them to detect trace gaps, errors, and changes after a rule is implemented. Remember that the declared agent can be spoofed; for sensitive analysis, verify IP addresses or mechanisms documented by the provider.<\/p>\n<p>Group by template and pattern, not by individual URL. Compare equivalent time periods and preserve the effect of caching. Reducing requests may be the goal, but confirm that important content is still being discovered and updated.<\/p>\n<h2 id=\"equipos\">Coordination between SEO, development and security<\/h2>\n<p>SEO defines which content needs crawling; development knows the routes and resources; security controls actual access. No team should use bots to achieve another&#039;s goal. Include the file in code review or an equivalent change process.<\/p>\n<p>When a campaign creates parameters or a platform changes routes, notify users before launch. Maintain a pattern dictionary with examples, purpose, responsible party, and expected rule. This discipline reduces accidental blocks and obsolete rules.<\/p>\n<p>Schedule a review after migrations and major updates. Compare the archived version with the public version and flag any unauthorized changes. If a vendor manages the archive, document the emergency channel and the expected remediation time; an after-hours incident should not depend on locating a single person.<\/p>\n<h2 id=\"checklist\">Checklist before saving<\/h2>\n<ul>\n<li>File in the root of the correct host.<\/li>\n<li>Public response and UTF-8 text.<\/li>\n<li>No blocking of strategic pages.<\/li>\n<li>Critical resources allowed.<\/li>\n<li>Correct objective: tracking, not security.<\/li>\n<li>absolute and canonical sitemap.<\/li>\n<li>Agent and employer testing.<\/li>\n<li>Copying, ownership, and monitoring.<\/li>\n<\/ul>\n<p>If you detect a critical conflict, don&#039;t add more rules without a map. You can <a href=\"https:\/\/www.seomos.com\/en\/contact\/\">contact SEOMOS<\/a> to review architecture and restoration.<\/p>\n<h2 id=\"fuentes\">Sources consulted<\/h2>\n<p>Documentation consulted on September 10, 2026.<\/p>\n<ul>\n<li><a href=\"https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/robots\/intro\" target=\"_blank\" rel=\"noopener\">Google: Introduction to robots.txt<\/a>.<\/li>\n<li><a href=\"https:\/\/developers.google.com\/crawling\/docs\/robots-txt\/create-robots-txt\" target=\"_blank\" rel=\"noopener\">Google: Create and test robots.txt<\/a>.<\/li>\n<li><a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9309\" target=\"_blank\" rel=\"noopener\">RFC 9309: Robots Exclusion Protocol<\/a>.<\/li>\n<\/ul>\n<h2 id=\"preguntas-frecuentes\">Frequently Asked Questions about robots.txt<\/h2>\n<details class=\"seomos-faq\">\n<summary class=\"seomos-faq-q\"><span>What does robots.txt do?<\/span><span class=\"seomos-faq-chev\" aria-hidden=\"true\">\u2304<\/span><\/summary>\n<div class=\"seomos-faq-a\">\n<p>It tells compatible crawlers which routes they can request on a host.<\/p>\n<\/div>\n<\/details>\n<details class=\"seomos-faq\">\n<summary class=\"seomos-faq-q\"><span>Does it prevent a page from appearing in Google?<\/span><span class=\"seomos-faq-chev\" aria-hidden=\"true\">\u2304<\/span><\/summary>\n<div class=\"seomos-faq-a\">\n<p>Not necessarily. A blocked URL can be indexed if it is discovered through links.<\/p>\n<\/div>\n<\/details>\n<details class=\"seomos-faq\">\n<summary class=\"seomos-faq-q\"><span>Does it protect private information?<\/span><span class=\"seomos-faq-chev\" aria-hidden=\"true\">\u2304<\/span><\/summary>\n<div class=\"seomos-faq-a\">\n<p>No. It uses authentication and authorization; the file is public and compliance is voluntary.<\/p>\n<\/div>\n<\/details>\n<details class=\"seomos-faq\">\n<summary class=\"seomos-faq-q\"><span>Where is it located?<\/span><span class=\"seomos-faq-chev\" aria-hidden=\"true\">\u2304<\/span><\/summary>\n<div class=\"seomos-faq-a\">\n<p>At the root of each protocol, the host and port to which it should be applied.<\/p>\n<\/div>\n<\/details>\n<details class=\"seomos-faq\">\n<summary class=\"seomos-faq-q\"><span>Should I include the sitemap?<\/span><span class=\"seomos-faq-chev\" aria-hidden=\"true\">\u2304<\/span><\/summary>\n<div class=\"seomos-faq-a\">\n<p>It can be declared with an absolute URL; it can also be submitted via Search Console.<\/p>\n<\/div>\n<\/details>\n<details class=\"seomos-faq\">\n<summary class=\"seomos-faq-q\"><span>How do I test for a change?<\/span><span class=\"seomos-faq-chev\" aria-hidden=\"true\">\u2304<\/span><\/summary>\n<div class=\"seomos-faq-a\">\n<p>Check the public archive and test real URLs per agent before and after deployment.<\/p>\n<\/div>\n<\/details>","protected":false},"excerpt":{"rendered":"<p>Robots.txt manages crawler access; it does not remove pages from the index or protect information. It learns syntax, cases, tests, and recovery from blocks.<\/p>","protected":false},"author":4,"featured_media":1339,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"content-type":"","footnotes":""},"categories":[11],"tags":[],"class_list":["post-1338","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-seo"],"_links":{"self":[{"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/posts\/1338","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/comments?post=1338"}],"version-history":[{"count":1,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/posts\/1338\/revisions"}],"predecessor-version":[{"id":1343,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/posts\/1338\/revisions\/1343"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/media\/1339"}],"wp:attachment":[{"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/media?parent=1338"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/categories?post=1338"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.seomos.com\/en\/wp-json\/wp\/v2\/tags?post=1338"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}