Most data APIs cache. It's the obvious move: fetch once, serve the same response to the next hundred callers, cut your own cost. Mindcase doesn't do this, and it's worth explaining exactly what that means in practice — not as a marketing line, but as an operational fact you can plan around.
What does "live on every run" actually mean?
It means what it says, literally: there's no cache layer sitting in front of the source. Every request Mindcase runs goes out and pulls the current state of the page at the moment you query it — not a snapshot from an hour ago, not a batch job that ran overnight.
This isn't a claim specific to one endpoint. It's the same language, independently, on the FAQ of every endpoint that addresses the question at all — 62 of the 72 endpoints in the real catalog have a "How fresh is the data?" entry, and all 62 say the same thing. A few real examples, verbatim:
- Amazon Products: "Live on every run. Amazon Products does not cache results. Each request pulls the current state of the source at the moment you query, so what you receive is what users see on the platform right now."
- LinkedIn Company Employees: "Live on every run. LinkedIn Company Employees does not cache results."
- Google Maps Places: "Live on every run. Results are never cached — each request pulls the current state of Google Maps at the moment you query, so you get exactly what users see right now."
- TikTok Posts: "Every request fetches live data directly from TikTok. There is no cache of stale results."
- Instagram Profiles: "Instagram Profiles fetches data live on every run, so you always get the current follower counts and bio info without any caching."
Different verticals, different teams' copy, same underlying behavior. This is a platform-wide architectural choice, not a per-endpoint setting.
Why does this matter more than it sounds?
Because the alternative — a cache — has a failure mode that's invisible until it isn't. A competitor drops a product's price on Monday and your tracker is still serving Friday's number. Someone gets promoted and your company-employees pull still lists their old title. A quick-commerce app runs out of stock and your feed says it's available.
None of those are edge cases. They're the entire reason to pull the data in the first place — you want to know what's true now, not what was true when a cache last refreshed. A cache doesn't just risk staleness; it makes staleness the default state between refreshes, and you don't find out which state you're in without checking twice.
What's the honest tradeoff?
A cache also does something useful: it absorbs the source's own problems. If the source site is briefly slow, a cache hit is still fast. If the source has a bad five minutes, a cache still serves the last good copy.
Mindcase doesn't have that shock absorber, and it's worth saying plainly rather than glossing over: a request is only as fast and as reliable as the live source is at that instant. There's no fallback copy to serve if the source is having a rough moment. That's the real cost of "always current" — you're exposed to the source's own availability, not insulated from it by a buffer.
For most of what people pull Mindcase for — prospecting lists, price checks, review audits — that tradeoff is the right one. For a use case that genuinely needs guaranteed uptime over guaranteed freshness, that's worth knowing before you build around it, not after.
How does this show up in a real response?
Freshness isn't just a policy statement — it's baked into how a request
actually resolves. A run call with wait=true executes the pull at
request time and returns the live rows in that same response:
POST /v1/data/linkedin/profiles/run?wait=true
If the job doesn't finish inside the wait window (a 60-second ceiling),
you get { job_id, status: "running" } back — the job keeps running
server-side, and you poll GET /jobs/{job_id}/results for the same live
pull to land. Nothing about polling makes the eventual result any less
live; it's the identical fetch, just returned asynchronously because it
took longer than the wait window. There's no separate "fast cached path"
and "slow live path" — one mechanism, always live, sometimes synchronous
and sometimes not depending on how long the source takes to respond.
Does every endpoint work this way?
Every endpoint whose FAQ speaks to it directly says yes — that's 62 of 72 in the current catalog, spanning e-commerce, social, local business, search, and travel verticals, with zero exceptions found. The remaining endpoints don't have that specific FAQ entry, which isn't evidence either way — check skills.md or an endpoint's own detail page for the current answer on any endpoint not listed above before assuming.
Which endpoint should you use for which job?
| Endpoint | Price | Best for |
|---|---|---|
| Amazon Products | $0.0015 / product | Price and listing monitoring where a stale price is worse than no price |
| Google Maps Places | $0.004 / place | Business listings where hours, ratings, and contact info change without notice |
| Instagram Profiles | $0.002 / profile | Follower counts and bio data that shift daily for active accounts |
| LinkedIn Company Employees | $0.004 / profile | Headcount and role data that's wrong the moment someone changes jobs |
| TikTok Posts | $0.002 / post | Engagement numbers that are meaningless if they're even an hour stale |
FAQ
No — it means the request always checks the current source, not that the source always changed. Pull the same product twice in a row and you'll get the same price if the price didn't move. The point is you're seeing the current state, not a cached one, whether or not that state happens to have changed since your last call.
No. Every run is a live pull, and pricing is per row returned regardless of whether the underlying value changed since your last request. If you want to avoid re-paying for unchanged data, that's a diffing step you build on your own side, not a caching mode Mindcase offers.
It means response time depends on the source, not on Mindcase. A source that responds quickly gets you a quick result; a slower source (or a large result set) is why the wait=true flow has a 60-second ceiling before falling back to polling. There's no cache to make a slow source feel fast.
The request fails rather than silently serving old data. You get an explicit failed status with an error, not a stale row presented as current — which is deliberate: a visibly failed request is easier to handle correctly than a quietly wrong one.
