Field notes on technical SEO diagnostics — crawl and indexation, audit method, and the mistakes that keep recurring across very different sites. the Amrut Services blog.
The writing published here is diagnostic rather than promotional. Each piece comes out of an actual engagement: a crawl that turned up something unexpected, a tool that reported a problem that was not a problem, a client question that turned out to be the wrong question. The intention is that a technical marketer or an in-house lead can read one and act on it without hiring anybody. Where a post connects to work described elsewhere on this site it is linked, but nothing here is a sales page with a headline on it.
Crawl bloat is the most common finding on large sites and the one clients are most surprised by, because nobody chose it — it accumulates. Faceted navigation, session parameters, tag archives, printer-friendly duplicates and staging leftovers all generate URLs that no human ever intended to exist, and Google indexes a great many of them. The post 9,220 pages indexed. Nobody decided to publish them. walks through finding the real ratio of intended to indexed pages, which is the single most useful number on a large site and the first thing measured in technical SEO.
A crawl report lists what a tool found. An audit says what it means, what it costs, and what to do about it in what order. The two get conflated constantly, usually by suppliers charging audit prices for crawl output. An SEO audit is not a crawl report sets out the difference and what a client should expect for the money — a prioritised, costed set of decisions rather than a two-hundred-row spreadsheet of severity flags. That distinction is the whole basis of how diagnostic work is scoped here.
A first crawl of a mid-sized site returns more issues than any team can act on, and most of them do not matter. The skill is not running the crawl, it is knowing which twenty rows out of forty thousand are worth an engineer's afternoon. How to read a crawl without drowning in it describes the passes to make and the order to make them in, so the output becomes a short list of decisions instead of a wall of red. It is the same filtering discipline applied on every engagement.
Prioritisation is where most SEO programmes fail, not analysis. Everything on a fix list is defensible in isolation; the question is which three items move revenue this quarter given the engineering capacity actually available. If you only do three things, which three? lays out how that call gets made — expected impact against implementation cost against the risk of doing nothing — and why a recommendation with no cost attached to it is not really a recommendation. The same framing shapes consulting engagements where somebody else does the building.
When pages fall out of the index the cause is almost always a change somebody made, and the fastest route to the answer is the deployment log rather than the SEO tooling. Deindexed pages: find out what shipped that week describes correlating the drop date against releases, and why "Google changed something" is the explanation to reach for last rather than first. This is the opening move in most recovery work, and it resolves a surprising share of cases inside a day.
A soft 404 report is easy to dismiss as a false positive, and it usually is not. If Google has decided a page is empty, thin or functionally a dead end, the honest reading is that a user would reach the same conclusion. Soft 404s: Google is usually right goes through the patterns that generate them — empty category pages, out-of-stock products, search results pages left indexable — and what to do with each. Most of the fixes are content and architecture decisions rather than technical ones, which is where on-page work takes over.
Crawl budget is the most over-applied concept in technical SEO. It is a genuine constraint on sites with hundreds of thousands of URLs and effectively irrelevant below that, yet it gets invoked to justify work on sites with four hundred pages. Crawl budget: almost nobody reading this has one explains where the threshold actually sits and what the real problem usually is when someone reaches for the term. Naming the wrong constraint is expensive, because the money then gets spent in the wrong place.
A missing robots.txt is not automatically a crisis, and a badly written one frequently is. The 1,630-page store with no robots.txt, and what that actually told me uses a real store to show what the absence of the file revealed about how the site had been built and maintained, and works through the disallow patterns that cause genuine damage. Robots handling sits alongside canonicals and sitemaps in the eligibility layer covered under technical SEO, and for stores specifically under ecommerce SEO.
1,562 missing H1s and seven H1s on one page. Same issue, opposite problem. takes two findings from the same crawl that look contradictory and shows they come from one underlying template decision. It is a worked example of why a severity count is not a diagnosis: the tool is reporting symptoms accurately and grouping them uselessly. Reading past the flag to the template that produced it is most of what separates an audit from a report.
Canonical tags are simple to describe and easy to get wrong in ways that suppress pages silently for months. Four ways canonical tags get misused, and one of them is my own advice covers the four patterns that recur, including one recommendation given here in the past that turned out to be wrong in a specific case. Publishing the correction rather than quietly dropping it is deliberate — anyone who has worked in this field for fifteen years has advice that aged badly, and pretending otherwise helps nobody.
Everything published here comes out of client work — mostly established sites with a diagnosis problem rather than a content problem. That work spans local SEO, Shopify and ecommerce stores, link acquisition, internal linking, off-page and white label delivery for agencies, and specialist verticals including criminal defence and personal injury firms, chiropractors and roofing and HVAC contractors. If a post raises a question about your own site, the contact page is the place to ask it.
Related SEO services: SEO Services · SEO Audit & Diagnostics · SEO Recovery · Local SEO · Ecommerce SEO · Shopify SEO · Technical SEO · On-Page SEO · Off-Page SEO · Link Building · Internal Linking Service · SEO Consulting · White Label SEO for Agencies.
Specialist work: Who I Work With · SEO for Criminal Defense Attorneys · SEO for Personal Injury Law Firms · SEO for Chiropractors · SEO for Roofers and HVAC Contractors.
More from Amrut Services: home · blog · about · contact · all Amrut Services documents · map.
Amrut Services — SEO consultant, Vadodara, Gujarat, India · +91 93769 10109 · mehuldedhiaa@gmail.com · amrutservices.com.