Skip to main content
BumpyLabs

How to Audit App Impact on Storefront Performance

A repeatable method for measuring what third-party Shopify apps actually cost your storefront, and a framework for deciding what to remove, defer, or replace.

Performance & CROThe BumpyLabs team13 min read

No merchant sets out to build a slow storefront. It happens one reasonable decision at a time.

A reviews app in year one, because social proof matters. A subscription app when the offer launched. A chat widget marketing asked for. An upsell tool someone trialled and never removed. Each one arrives with a clear rationale and a small measured cost. None of them is the problem. Collectively they are.

The reason this goes unnoticed is that nobody owns the total. Individual apps are evaluated on what they do, almost never on what they cost, and the cost compounds quietly in a place — the storefront's critical rendering path — that no dashboard reports on.

What "slow" actually means

Before measuring anything, agree what you are measuring against. Google's Core Web Vitals are the practical standard, because they are what search actually assesses.

MetricWhat it measuresGoodNeeds improvementPoor
LCP (Largest Contentful Paint)How long until the main content renders≤ 2.5s2.5s–4.0s> 4.0s
INP (Interaction to Next Paint)How quickly the page responds to input≤ 200ms200ms–500ms> 500ms
CLS (Cumulative Layout Shift)How much the layout moves unexpectedly≤ 0.10.1–0.25> 0.25

Two details matter for how you interpret these.

INP replaced First Input Delay in March 2024, and FID was removed from Chrome's tooling entirely later that year. If a report or an agency proposal still references FID, it is out of date. INP is stricter in a useful way: it measures the full delay from interaction to visible response, across the whole page visit, rather than just the first interaction.

Assessment is at the 75th percentile of real users. Not the average. This is why a store that feels fine on a developer's laptop can fail — the measurement is weighted toward the slower quarter of your actual traffic, on real phones and real networks.

Third-party apps map onto these metrics predictably. Render-blocking scripts hurt LCP. Heavy JavaScript execution hurts INP. Widgets that inject content after load — banners, badges, popups — hurt CLS.

Field data versus lab data

This distinction causes more confusion than anything else in performance work, and getting it wrong leads teams to optimize for the wrong signal.

Field data is measured from real Chrome users on your site, collected in the Chrome User Experience Report and surfaced in Search Console and PageSpeed Insights. It reflects genuine devices, connections, and behaviour. It is what Google uses.

Lab data is a simulated test run under fixed conditions — Lighthouse, WebPageTest, the Chrome DevTools performance panel. It is reproducible, which makes it excellent for diagnosis, and artificial, which makes it unreliable as a verdict.

They disagree routinely, and both can be right. A store might pass in the field because most traffic is on good connections in one city, while Lighthouse's throttled mobile simulation fails it. Or lab tests look clean while field data is poor, because real users trigger app behaviour that a cold synthetic load never reaches.

Use field data to decide whether you have a problem. Use lab data to find out why.

Build the app inventory

You cannot audit what you have not listed. Start with the admin, then verify against the storefront, because the two disagree more often than you would expect.

For each installed app, record:

  • What it does, and which business function relies on it
  • Monthly cost
  • How it integrates — theme app extension, script tag, or manually installed snippet
  • Which pages it affects
  • When it was last evaluated, and by whom

Then check what the storefront actually loads, because the admin list is incomplete. Apps removed in the past can leave snippets behind, and manually installed code frequently outlives the app it belonged to.

Open a product page with DevTools, and look for third-party origins in the network panel that you cannot attribute to anything in your app list. Every unattributed request is a finding.

Measure what each app costs

The goal is attribution: not "the storefront is slow" but "this app costs this much on this page".

Establish a baseline

Test the same representative pages every time — typically the homepage, a busy collection page, and a product page, since they load different things. Run each test several times and take the median. Single runs vary enough to mislead.

Record LCP, INP or a total blocking time proxy, CLS, total JavaScript transferred, and total requests.

Attribute cost by origin

In DevTools' network panel, group requests by domain. Third-party app scripts are usually easy to identify by hostname. For each third-party origin, note transfer size, how long the main thread spends executing it, and — most importantly — whether it loads before or after first render.

That last point is where the biggest wins hide. A 30KB script that blocks rendering costs more than a 200KB one that loads afterwards. Size is the number everyone quotes; position in the critical path is the number that matters.

Test with the app disabled

This is the strongest evidence available, and it is what turns an argument into a decision.

Duplicate your live theme so you have a safe sandbox. In the copy, disable one app's storefront output — usually by turning off its theme app extension block or removing its snippet. Re-run the same tests on the same pages. The difference is that app's real cost on your store, not a vendor's benchmark.

Repeat for each significant app. It is methodical rather than difficult, and it produces a table nobody can argue with.

Check what each app loads where

Many apps load everywhere by default, whether or not they are needed. A reviews widget belongs on product pages, not on the cart. A subscription script is irrelevant on your policy pages. Check each app's page-level footprint — apps loading on pages where they serve no purpose are the easiest wins available, and they usually require configuration rather than removal.

Reading the results

You now have a cost per app. Sort by cost and look for the shape of the problem.

Usually one or two apps dominate — often ones nobody has questioned in years because they were installed before anyone currently on the team joined. Occasionally there is no single culprit and the total is death by a dozen small scripts, which is a harder conversation because no individual removal moves the number much.

Set the findings against value. An app costing 400ms of LCP that drives a meaningful share of revenue is a different proposition from one costing 300ms that displays a badge nobody clicks. The performance number alone does not decide anything; it just makes the trade-off visible so the business can weigh it.

What different app categories typically cost

App categories behave differently, and knowing the usual pattern tells you where to look first. These are tendencies rather than rules — a well-built app in a heavy category can easily outperform a careless one in a light category.

App categoryTypical footprintUsual metric affectedWhere the cost hides
Reviews and ratingsHeavyLCP, CLSLoads on every page; injects content after render, moving layout
Chat and support widgetsHeavyINP, LCPLarge bundles loading immediately when they could wait for intent
Upsell and cross-sellModerate to heavyINP, CLSRuns on product and cart; often inserts elements late
SubscriptionsModerateLCPProduct-page logic that frequently loads site-wide
Search and filteringModerateINPJustifiable where used; wasteful when loaded on every template
Analytics and pixelsLight each, heavy togetherINPIndividually small, collectively significant; rarely audited
Banners and popupsLight to moderateCLSInjecting above existing content is a common layout-shift cause

The pattern worth noting: several categories load site-wide when they are only needed on one template. That is usually a configuration fix rather than a removal, which makes it the cheapest kind of win available.

A worked example of the method

To make the sequence concrete, here is how a product-page audit typically runs.

Baseline. Test the product page five times, take the median, and record it: LCP, INP, CLS, total JavaScript, and request count. This is the number every later comparison refers back to.

Attribution. Group network requests by origin. Say four third-party domains account for most of the JavaScript — reviews, chat, an upsell tool, and an analytics vendor. Note the transfer size, main-thread execution time, and whether each loads before first render.

Isolation. Duplicate the theme. Disable the reviews app's storefront output. Re-run the same five tests. The delta is that app's real cost on this page. Restore it, then repeat for each of the other three.

Judgement. Now the conversation changes. Instead of "the site feels slow", you have four numbers next to four business functions, and the question becomes whether each function is worth its cost. Reviews on a product page usually are. A chat widget loading eagerly on a page where almost nobody opens it usually is not — and that one is a defer, not a removal.

Verification. Make one change, re-measure, record the result. Then the next.

The whole loop is a few hours of methodical work. Its value is that it converts an argument about priorities into a table of measurements.

Setting a performance budget

Auditing once fixes today. A budget is what stops the problem returning.

A budget is a number the storefront is not allowed to exceed without a deliberate decision — a JavaScript weight ceiling for key templates, a maximum count of third-party origins, or a target for LCP on product pages. It does not need tooling to start. It needs agreement, and a check before anything new goes on.

Make it a question rather than a process: what does this load, on which pages, and does the page wait for it? Anyone can ask that, and asking it before installation is what prevents the slow accumulation that makes audits necessary in the first place.

Remove, defer, or replace

Every finding resolves into one of three actions, with different risk profiles.

Remove. The cleanest outcome, appropriate when the app is unused, duplicated by another tool, or delivering less than it costs. Confirm nobody depends on it, then remove it — and afterwards verify the storefront no longer requests it. Leftover snippets are extremely common.

Defer. Appropriate when the function matters but does not need to be available at first render. Chat widgets, review carousels below the fold, and social embeds usually fall here. The function is preserved; the cost moves out of the critical path. Some apps expose loading options; others need theme-level work.

Replace. Appropriate when the function matters and the app is the wrong implementation of it. Sometimes the replacement is a lighter competitor. Sometimes it is native Shopify functionality that has caught up since the app was installed. Sometimes it is a small piece of custom code that does one thing well instead of a general-purpose tool doing forty.

Whichever route, change one thing at a time and re-measure. Changing five things and observing a 20% improvement teaches you nothing about which change to repeat.

What to do after

Performance is a state you maintain, not a project you finish.

Re-measure after each change, on the same pages under the same conditions, and keep the before-and-after. This is the record that justifies the work and settles the question of whether it helped.

Watch field data over the following weeks. CrUX reports on a rolling window, so improvements take time to appear in Search Console. A change that looks good in the lab needs a few weeks to confirm in the field.

Put a gate on new app installs. Not a bureaucratic process — just one question before installing anything: what does this load, where, and does the page wait for it? Trialling an app on a duplicated theme first costs an hour and prevents most of the accumulation this article is about.

Schedule a review. Twice a year, walk the app list and ask what each one is still doing. Apps are easy to install and easy to forget, which is exactly how storefronts get slow one reasonable decision at a time.

The method, condensed

List every app and verify against what the storefront actually requests. Baseline your key pages with both field and lab data. Attribute cost per app by testing with each one disabled on a duplicated theme. Sort by cost, weigh against value, then remove, defer, or replace — one change at a time, re-measuring after each.

None of it is complicated. It is just work that nobody owns until somebody decides to.

If you want this run on your store with the findings written up as a prioritized backlog rather than a list of scores, that is a service we offer — and we will tell you honestly if the answer is that your apps are fine and the problem is somewhere else.

FAQ

Questions people ask us about this

There is no useful number. Ten well-behaved apps that load only where they are needed can cost less than three that inject scripts on every page. Judge apps by what they load, where they load it, and whether the storefront waits for them — not by the count in your admin.

Keep reading

Related insights

  • Migration

    How to Plan a Migration to Shopify

    A practical guide to planning a Shopify migration: what to audit before you scope, how to map fields and URLs, and where migrations usually go wrong.

    12 min readRead
  • Apps & Integrations

    When a Custom Shopify App Is Worth Building

    A practical test for choosing between an app from the Shopify App Store, a workflow automation tool, and custom development — and what a custom app really costs to own.

    11 min readRead

Next step

Want this handled by a Shopify-only team?

Share where your store is today and the constraint you are working around. We will tell you what we would actually do next.