Identity
- Handle, display name, profile URL
- Linked accounts across platforms
- Account type and category
- Language, country and city, where public
Clean, true, fast
Ask five companies what creator data is and you’ll get five answers. We will share our pillars and why we believe we are building the best infrastructure in the market.
Back in 2023 I scrapped everything on our product roadmap except the API. CRM view, outreach module, workflows, integrations, all of it.
We were never going to win the features game. But we were going to win the data game, because that’s what we’re good at.
Three years later that bet has a name. We’re creator data infrastructure, the data layer underneath the products that discover, vet, onboard and pay creators. Most of the time you don’t see us. You see a creator marketplace, an onboarding flow, an affiliate program, and somewhere under the hood there’s a call to our API.
That said, “creator data” has become one of those phrases that means whatever the person saying it needs it to mean. A follower count, a guessed email and a demographic chart with no method behind it all get called creator data.
From the outside, a buyer can’t tell the difference, and honestly, a lot of the time the data provider can’t explain it from the inside either.
So this is me being extremely straightforward about what we have, what we don’t have, what we’re building, and how we check it.
My ex-boss used to say, “Nikola, people always remember 3 main things,” and that’s true now.
There are three things a data provider has to get right. The data has to be clean, true and fast.
Where it comes from, what we refuse to hold, and the legal side.
How close it is to the real creator. Accuracy and coverage per data point, published.
How fast you can build on it. Monthly plans, an API key on signup, docs, MCP, playbooks.
Before those three, the vocabulary. When we say creator data, this is the record we mean.
Every record starts with identity: the handle, the display name, the linked accounts across platforms. That last part matters.
A lot of what makes a record useful is knowing which accounts belong to the same human, and we wrote up how that works in its own post.
After identity comes reach and engagement: followers, the historical follower series, growth, how often they post, what their recent posts actually did, and how sponsored posts performed against the rest.
The audience block has top markets, languages, age and gender distribution and niche interests. And again that last one matters because so far, niche has been really broad.
Audience interested in “Music”. Sure, but what type of music? Pop, rap, christian rock? It definitely matters.
Contact is the publicly listed contact info, its verification status, link-in-bio destinations and the personal site.
And the commercial signals are monetization status, brands mentioned and promoted, affiliate and creator-program participation, and the topics and keywords they use.
When I cut the roadmap in 2023, I said the only three things worth betting on were contact data, the verified social graph and historical insights, because those were the hard things to copy. Every field on this record exists to serve one of those three.
If you’re going to build a product on someone else’s data, you need three things from them: where it comes from, what they won’t touch, and a legal side your team can sign off on.
Three sources, combined into one record, and all of it is publicly available information.
The public internet: search results, personal sites, press, public directories and creator listings. Third-party partners who supply publicly available data, mostly contact info, filling gaps where we don’t already hold it ourselves. And our own research team, whose entire job is the state of the database. They find and fix errors in the pipeline, benchmark the quality numbers, and vet creators every day.
There’s no fourth source. And because we own 100% of our data, when a customer asks “where did this field come from”, we can answer.
In this category that isn’t a given, and it’s the single thing I’d check first if I were buying.
The clearest way to describe a dataset is to draw its edges.
Nothing behind a password or a login wall, no private messages or non-public account activity, no payment, financial or government identifiers, no special-category personal data, and no records of creators who have opted out.
I want to be super direct here, because this is the part that costs us deals, and I’d rather say it here than on a sales call.
If it isn’t publicly available, we don’t have it. That means no story data, no shares and saves, and no sales volume for a creator’s shop. We get asked for all three, and the answer is the same every time: we can model some of it from public signals, and we won’t go behind a login to get the raw number.
A standard DPA is available to every customer, with a published subprocessor list, so your legal team can review the chain before you sign rather than after. Any creator can ask what we hold about them, correct it, or have it removed, and removals hold on later refreshes. Our entities, policies and legal contacts are published in one place instead of scattered through footers.
Enterprise buyers don’t want a paragraph here. They want a page they can send to their legal team, and that page exists.
True means the record matches the real creator today, not the day we first found them.
Every creator database decays. Handles change, accounts go quiet, emails stop working, and a creator who was in Lisbon last year is in Bali this year. So for us, this is a process with numbers attached, not a launch feature.
The question I keep asking our data team is the one a customer would ask: how do we know location is accurate, what’s the methodology, and what’s the chance of error?
So for every data point we roll out, we try to publish three things:
Every provider says their data is accurate. We’re the first to show our work. The method, the accuracy and the threshold for every data point are on our data methodology page.
Here’s where we are today.
Two things here. None of the accuracy numbers are 100%, and none of them ever will be, because I’ve never seen a data point that is. And the numbers move. Location was the hardest one for us.
Accuracy is one number. Coverage is the other, and it’s a different question: of the records we hold, how many carry this field at all? Some fields are close to universal, some aren’t, and the page says which.
We don’t trust a single source for an accuracy number, including our own pipeline. AI on its own can be biased by what it was trained on, and human reviewers on their own can’t cover the scale. So every number comes out of two tracks run in parallel, and then compared.
On the first track, at least two independent AI sources label the same data, tens of thousands of records per cycle, and the disagreements get surfaced rather than averaged away.
On the second, a stratified manual review: at least three thousand records per subset, with subsets cut across country, language, follower bucket, account type and niche, and reviewers who can’t see what the pipeline predicted. Two AI sources agreeing is a much stronger signal than one. A human catching the edge case that random sampling would miss is stronger still.
With the method and the confidence threshold next to it.
Every profile is revisited at least monthly, and accounts that change quickly get picked up more often.
On top of that, anything you export is refreshed at the point of export, so what lands in your CSV is current on the day you use it, not the day we last ran the cycle.
Audience data is where this space is a bit all over the place, so here’s exactly how ours works.
We analyse the public engagers on a creator’s content and classify each one the same way we classify a creator: from the public bio, the category, the hashtags and the recent captions. That gives a two-level picture, 32 interests at the top and 413 sub-interests underneath.
So a share is a count of people, not a guess. When a creator’s audience reads 19.8% music, one in five of the people engaging with them post about music themselves.
This is the part I’d push on if you’re comparing us to anyone. Most audience data is read off who an audience follows. Ours is read off what the audience actually posts. And because we hold the engager profiles ourselves, anything we can build at the creator level becomes an audience-level lens automatically, which is not something you can do while reselling somebody else’s report.
Fifteen audience fields ship per creator: gender, age, geography, language, interests, sub-interests, credibility and reachability. All of it is derived from publicly available data, it’s labelled as an estimate in the product, and it’s live across the major platforms today with more coverage landing every month.
It’s a model of the audience, not a platform readout, and we say so in the product.
This one gets the strictest treatment because it’s the one that costs you money when it’s wrong.
Finding an email is easy. Creators publish them in their public bios, in about sections, on their own sites. Identifying the one a brand should trust is the hard part, and a single profile can surface several: a real partnership inbox, a typo, a role-based mailbox, a spam-pattern string that happens to look like an address.
Every candidate goes through format and syntax checks before it’s stored. Typos we recognize get repaired, and anything malformed past repair is dropped. Then each address is scored by how much we trust it. An email on a domain the creator owns scores highest. A role-based or spam-pattern inbox scores low. And at export, every email is checked again across multiple independent verification providers, so one provider’s blind spot doesn’t become your bounce. If it fails, we don’t sell it. That part is real-time.
Bios, about sections, contact pages, the creator’s own site. The same address across several places raises confidence.
Every candidate is checked before it’s stored. Recognisable typos are repaired. Malformed past repair is dropped.
Creator-owned domains score highest. Role-based and spam-pattern inboxes score low. The best address wins.
Checked across multiple independent providers at the moment you export. If it fails, it isn’t sold.
It’s very hard for anybody to have all profiles. We certainly try, and it’s still very hard. A creator who posted twice in 2019 and vanished might not be in our index, and a famous account with a handle nobody links to can resolve later than you’d expect. Estimates are labelled as estimates in the product, and when we don’t have a field it stays blank.
There’s also no sandbox to test against. We give you credits on the real data instead, because a fake dataset tells you nothing about whether ours is any good.
If it takes six weeks and three calls to get access, it isn’t infrastructure yet.
Fast, for us, is one number: how fast can someone go from “I need creator data” to a working call, without talking to a human? Everything below exists to make that number small.
Start on a monthly plan and cancel any time. Credits roll over, and seats are unlimited on every plan, so putting your whole team on it costs nothing extra. I’d rather earn the renewal every month than lock it into a contract. That said, if you want an annual plan we do those too. But you never have to start there.
Starter is $140 a month billed annually, or $199 month to month. Pro starts at $208. The price per credit is published too, as low as $0.15. You can work out what we cost before you talk to anyone, which is not how most of this category sells data.
Sign up and your API key is on the page, with ten free credits to test with and the docs at docs.influencers.club. No payment details to start. It’s the same dataset behind the dashboard, the API, our MCP server and custom exports.
Native connections to HubSpot, Salesforce and Snowflake, and nodes for Clay, Zapier and n8n. The n8n node exists because every one of our playbooks used to assume you had engineers to build the integration, and now you don’t need them.
influencers.club is an official ChatGPT connector. One click, sign in with your account, and ChatGPT can search and enrich creators like any other tool it has. On claude.ai it’s a hosted connector, and there’s an open-source local server for Cursor, VS Code, Claude Desktop, Windsurf and Zed.
Plug it in and just ask. The answer comes from the actual records, not from a model’s guess, and setup takes about two minutes.
There are sixteen of them at last count. Creatorbooks are our way of documenting the best, most clever things companies have actually done using creator data in their real workflows, like finding the creators hiding in your e-commerce customers, screening applications with fake follower checks, or personalizing onboarding when a signup turns out to be a creator. Each one is step by step, with the API calls.
We run shared Slack channels with a lot of the teams building on us, so when something looks off, their engineers talk to our engineers directly instead of going through a ticket.
And the proof it works: Kajabi put creator enrichment into their onboarding flow and lifted onboarding conversions by 10%. Pearpop onboarded 100K+ creators on our data. Both are written up with the implementation, not just the number.
Search and filter in the browser, plus natural-language search. Export what you find.
Discovery, enrichment and lookup endpoints, used in production onboarding flows and marketplaces.
Official ChatGPT connector, hosted on claude.ai, open source everywhere else. Ask in plain language.
Custom datasets and reports, one-off or on a schedule.
Everything above is checkable. The accuracy numbers live on data.influencers.club with the method next to each one. The sourcing, the edges of the dataset, creator rights and the DPA live on influencers.club/our-data.
When a number changes, the page changes. When we add a data point, it goes on that page with its method.
Every number in this post is published there. Go check.