Clean, true, fast

Our promise: Creator data infrastructure, built clean, true, and fast.

Ask five companies what creator data is and you’ll get five answers. We will share our pillars and why we believe we are building the best infrastructure in the market.

Back in 2023 I scrapped everything on our product roadmap except the API. CRM view, outreach module, workflows, integrations, all of it.

We were never going to win the features game. But we were going to win the data game, because that’s what we’re good at.

Three years later that bet has a name. We’re creator data infrastructure, the data layer underneath the products that discover, vet, onboard and pay creators. Most of the time you don’t see us. You see a creator marketplace, an onboarding flow, an affiliate program, and somewhere under the hood there’s a call to our API.

That said, “creator data” has become one of those phrases that means whatever the person saying it needs it to mean. A follower count, a guessed email and a demographic chart with no method behind it all get called creator data.

From the outside, a buyer can’t tell the difference, and honestly, a lot of the time the data provider can’t explain it from the inside either.

So this is me being extremely straightforward about what we have, what we don’t have, what we’re building, and how we check it.

My ex-boss used to say, “Nikola, people always remember 3 main things,” and that’s true now.

There are three things a data provider has to get right. The data has to be clean, true and fast.

01

Clean

Where it comes from, what we refuse to hold, and the legal side.

02

True

How close it is to the real creator. Accuracy and coverage per data point, published.

03

Fast

How fast you can build on it. Monthly plans, an API key on signup, docs, MCP, playbooks.

First, what a creator record is

Before those three, the vocabulary. When we say creator data, this is the record we mean.

Every record starts with identity: the handle, the display name, the linked accounts across platforms. That last part matters.

A lot of what makes a record useful is knowing which accounts belong to the same human, and we wrote up how that works in its own post.

After identity comes reach and engagement: followers, the historical follower series, growth, how often they post, what their recent posts actually did, and how sponsored posts performed against the rest.

The audience block has top markets, languages, age and gender distribution and niche interests. And again that last one matters because so far, niche has been really broad.

Audience interested in “Music”. Sure, but what type of music? Pop, rap, christian rock? It definitely matters.

Contact is the publicly listed contact info, its verification status, link-in-bio destinations and the personal site.

And the commercial signals are monetization status, brands mentioned and promoted, affiliate and creator-program participation, and the topics and keywords they use.

When I cut the roadmap in 2023, I said the only three things worth betting on were contact data, the verified social graph and historical insights, because those were the hard things to copy. Every field on this record exists to serve one of those three.

What a creator record holdsFIG. 01

Identity

  • Handle, display name, profile URL
  • Linked accounts across platforms
  • Account type and category
  • Language, country and city, where public

Reach

  • Followers and subscribers
  • Historical follower series
  • Growth over 30 / 90 days
  • Content volume and cadence

Engagement

  • Average likes and comments
  • Engagement rate
  • Recent post performance
  • Performance on sponsored posts

Audience

  • Top markets and languages
  • Age and gender distribution
  • Audience interests
  • Niches and subniches

Contact

  • Publicly listed contact info
  • Verification status, multi-provider
  • Link-in-bio destinations
  • Personal site and other socials

Commercial signals

  • Monetization status
  • Brands mentioned and promoted
  • Affiliate and creator-program participation
  • Keywords and topics used
The fields most teams work with day to day. A record carries more beyond them. Source: influencers.club/our-data

Clean: where it comes from, and where it stops

If you’re going to build a product on someone else’s data, you need three things from them: where it comes from, what they won’t touch, and a legal side your team can sign off on.

Where it comes from

Three sources, combined into one record, and all of it is publicly available information.

The public internet: search results, personal sites, press, public directories and creator listings. Third-party partners who supply publicly available data, mostly contact info, filling gaps where we don’t already hold it ourselves. And our own research team, whose entire job is the state of the database. They find and fix errors in the pipeline, benchmark the quality numbers, and vet creators every day.

There’s no fourth source. And because we own 100% of our data, when a customer asks “where did this field come from”, we can answer.

In this category that isn’t a given, and it’s the single thing I’d check first if I were buying.

What we refuse to hold

The clearest way to describe a dataset is to draw its edges.

Nothing behind a password or a login wall, no private messages or non-public account activity, no payment, financial or government identifiers, no special-category personal data, and no records of creators who have opted out.

I want to be super direct here, because this is the part that costs us deals, and I’d rather say it here than on a sales call.

If it isn’t publicly available, we don’t have it. That means no story data, no shares and saves, and no sales volume for a creator’s shop. We get asked for all three, and the answer is the same every time: we can model some of it from public signals, and we won’t go behind a login to get the raw number.

The legal side

A standard DPA is available to every customer, with a published subprocessor list, so your legal team can review the chain before you sign rather than after. Any creator can ask what we hold about them, correct it, or have it removed, and removals hold on later refreshes. Our entities, policies and legal contacts are published in one place instead of scattered through footers.

Enterprise buyers don’t want a paragraph here. They want a page they can send to their legal team, and that page exists.

The edges of the datasetFIG. 02

In the dataset

  • Publicly visible profile and account information
  • Public content, captions, keywords and links
  • Contact details creators publish for business enquiries
  • Public engagement, and audience estimates modelled from it
  • Publicly available data supplied by partners

×Not in the dataset

  • Anything behind a password or login wall
  • Private messages or non-public account activity
  • Payment, financial or government identifiers
  • Special-category personal data
  • Records of creators who have opted out
Our data policy, as published. Source: influencers.club/our-data

True: accuracy we publish, not accuracy we promise

True means the record matches the real creator today, not the day we first found them.

Every creator database decays. Handles change, accounts go quiet, emails stop working, and a creator who was in Lisbon last year is in Bali this year. So for us, this is a process with numbers attached, not a launch feature.

The question I keep asking our data team is the one a customer would ask: how do we know location is accurate, what’s the methodology, and what’s the chance of error?

So for every data point we roll out, we try to publish three things:

  • The average accuracy we’re comfortable guaranteeing for it
  • a short description of how we get there
  • the confidence threshold we use.

Every provider says their data is accurate. We’re the first to show our work. The method, the accuracy and the threshold for every data point are on our data methodology page.

Here’s where we are today.

Published accuracy, by data pointFIG. 03
Bar chart 0% 25% 50% 75% 100% First name: 96% First name 96% Profile matching: 96% Profile matching 96% Location: 95% · avg Location 95% avg Type of profile: 93% · avg F1 Type of profile 93% avg F1 Language: 93% Language 93% Gender: 92% · avg Gender 92% avg Audience analytics: 90% Audience analytics 90%
Accuracy against reviewed ground truth, per data point. Type of profile is reported as F1; location and gender as averages across subsets. Source: data.influencers.club, September 2026

Two things here. None of the accuracy numbers are 100%, and none of them ever will be, because I’ve never seen a data point that is. And the numbers move. Location was the hardest one for us.

Accuracy is one number. Coverage is the other, and it’s a different question: of the records we hold, how many carry this field at all? Some fields are close to universal, some aren’t, and the page says which.

Published coverage, by data pointFIG. 04
Bar chart 0% 25% 50% 75% 100% Engagement rate: 100% Engagement rate 100% Follower growth: 100% Follower growth 100% Posting frequency: ~95% Posting frequency ~95% Hashtags & mentions: ~90% Hashtags & mentions ~90% Views growth: ~80% Views growth ~80% Links in bio: ~72% Links in bio ~72%
Share of records that carry the field. Approximate values are marked with a tilde on the methodology page and here. Source: data.influencers.club, September 2026

How the numbers get made

We don’t trust a single source for an accuracy number, including our own pipeline. AI on its own can be biased by what it was trained on, and human reviewers on their own can’t cover the scale. So every number comes out of two tracks run in parallel, and then compared.

On the first track, at least two independent AI sources label the same data, tens of thousands of records per cycle, and the disagreements get surfaced rather than averaged away.

On the second, a stratified manual review: at least three thousand records per subset, with subsets cut across country, language, follower bucket, account type and niche, and reviewers who can’t see what the pipeline predicted. Two AI sources agreeing is a much stronger signal than one. A human catching the edge case that random sampling would miss is stronger still.

Two tracks, one accuracy numberFIG. 05

Multi-AI cross-validation

TRACK A
  • At least 2 independent AI sources label the same data
  • Tens of thousands of records per cycle
  • Disagreements surfaced, never averaged away

Stratified manual review

TRACK B
  • At least 1,000 records reviewed per subset
  • Subsets across country, language, follower bucket, account type, niche
  • Reviewers blind to the pipeline’s prediction

Compare

BOTH TRACKS
  • Agreement between tracks sets the number
  • Disagreement becomes the next fix
  • Re-run every validation cycle
PUBLISHED
Accuracy per data point

With the method and the confidence threshold next to it.

The validation process behind every accuracy number on the methodology page. Source: data.influencers.club

Freshness

Every profile is revisited at least monthly, and accounts that change quickly get picked up more often.

On top of that, anything you export is refreshed at the point of export, so what lands in your CSV is current on the day you use it, not the day we last ran the cycle.

Audience data

Audience data is where this space is a bit all over the place, so here’s exactly how ours works.

We analyse the public engagers on a creator’s content and classify each one the same way we classify a creator: from the public bio, the category, the hashtags and the recent captions. That gives a two-level picture, 32 interests at the top and 413 sub-interests underneath.

How deep the audience picture goesFIG. 06
32audience interests at the top level of the taxonomy
413sub-interests underneath, read from what the audience posts
15audience fields per creator
Sub-interests per audience interest, all 32 Arts & Design: 22 sub-interests Arts & Design 22 Music: 20 sub-interests Music 20 Entertainment & Media: 19 sub-interests Entertainment & Media 19 Fashion, Apparel & Accessories: 19 sub-interests Fashion, Apparel & Accessories 19 Business & Professional Services: 18 sub-interests Business & Professional Services 18 Education & Learning: 18 sub-interests Education & Learning 18 Automotive & Transportation: 17 sub-interests Automotive & Transportation 17 Community & Nonprofits: 16 sub-interests Community & Nonprofits 16 Technology & Software: 16 sub-interests Technology & Software 16 Fitness & Exercise: 15 sub-interests Fitness & Exercise 15 Sports: 15 sub-interests Sports 15 Travel & Tourism: 15 sub-interests Travel & Tourism 15 Food & Nutrition: 14 sub-interests Food & Nutrition 14 Beauty & Grooming: 13 sub-interests Beauty & Grooming 13 Beverages & Drinks: 12 sub-interests Beverages & Drinks 12 Comedy & Memes: 12 sub-interests Comedy & Memes 12 Finance & Investing: 12 sub-interests Finance & Investing 12 Lifestyle & Personal Development: 12 sub-interests Lifestyle & Personal Development 12 Religion & Spirituality: 12 sub-interests Religion & Spirituality 12 Legal & Politics: 11 sub-interests Legal & Politics 11 Outdoor & Adventure: 11 sub-interests Outdoor & Adventure 11 Health & Wellness: 10 sub-interests Health & Wellness 10 Photography & Videography: 10 sub-interests Photography & Videography 10 Adult Content: 9 sub-interests Adult Content 9 Home & Living: 9 sub-interests Home & Living 9 News & Commentary: 9 sub-interests News & Commentary 9 Pets & Animals: 9 sub-interests Pets & Animals 9 Retail & Shopping: 9 sub-interests Retail & Shopping 9 Events & Weddings: 8 sub-interests Events & Weddings 8 Gaming & Esports: 8 sub-interests Gaming & Esports 8 Kids & Parenting: 8 sub-interests Kids & Parenting 8 Real Estate: 5 sub-interests Real Estate 5
Music, one of the 32, opened up: 20 sub-interests underneath it.
Pop MusicRock MusicHip HopJazz MusicClassical MusicR&B SoulCountry MusicReggae MusicBlues MusicMetal MusicIndie MusicHouse MusicWorld & Folk MusicMusic ProductionInstrumentsVocal ArtsMusic EducationMusic Fandom & CultureMusic Technology & GearReligious & Spiritual Music
All 32 interests, ranked by how many sub-interests sit under each. Every engager is scored against all 413. Source: influencers.club taxonomy v6, the version the classifier runs today

So a share is a count of people, not a guess. When a creator’s audience reads 19.8% music, one in five of the people engaging with them post about music themselves.

This is the part I’d push on if you’re comparing us to anyone. Most audience data is read off who an audience follows. Ours is read off what the audience actually posts. And because we hold the engager profiles ourselves, anything we can build at the creator level becomes an audience-level lens automatically, which is not something you can do while reselling somebody else’s report.

Fifteen audience fields ship per creator: gender, age, geography, language, interests, sub-interests, credibility and reachability. All of it is derived from publicly available data, it’s labelled as an estimate in the product, and it’s live across the major platforms today with more coverage landing every month.

It’s a model of the audience, not a platform readout, and we say so in the product.

Validated accuracy on audience fieldsFIG. 07
Validated accuracy on audience fields 0% 25% 50% 75% 100% Gender: 97% Gender 97% Niche analysis: 95% Niche analysis 95% Location: 94% Location 94% Audience interests: 92% Audience interests 92%
Validated accuracy per audience field, measured on 1,000+ profiles per use case. Definition, sources, algorithm and validation results per data point live on the methodology page. Source: data.influencers.club

Email, the field people build outreach on

This one gets the strictest treatment because it’s the one that costs you money when it’s wrong.

Finding an email is easy. Creators publish them in their public bios, in about sections, on their own sites. Identifying the one a brand should trust is the hard part, and a single profile can surface several: a real partnership inbox, a typo, a role-based mailbox, a spam-pattern string that happens to look like an address.

Every candidate goes through format and syntax checks before it’s stored. Typos we recognize get repaired, and anything malformed past repair is dropped. Then each address is scored by how much we trust it. An email on a domain the creator owns scores highest. A role-based or spam-pattern inbox scores low. And at export, every email is checked again across multiple independent verification providers, so one provider’s blind spot doesn’t become your bounce. If it fails, we don’t sell it. That part is real-time.

How an email earns its place on a recordFIG. 08
STEP 1

Found where creators publish it

Bios, about sections, contact pages, the creator’s own site. The same address across several places raises confidence.

STEP 2

Format and syntax checks

Every candidate is checked before it’s stored. Recognisable typos are repaired. Malformed past repair is dropped.

STEP 3

Scored by trust

Creator-owned domains score highest. Role-based and spam-pattern inboxes score low. The best address wins.

STEP 4

Verified again at export

Checked across multiple independent providers at the moment you export. If it fails, it isn’t sold.

Example candidates from the methodology page. Green passes, struck-through is repaired or dropped. Source: data.influencers.club/email

What we don’t have

It’s very hard for anybody to have all profiles. We certainly try, and it’s still very hard. A creator who posted twice in 2019 and vanished might not be in our index, and a famous account with a handle nobody links to can resolve later than you’d expect. Estimates are labelled as estimates in the product, and when we don’t have a field it stays blank.

There’s also no sandbox to test against. We give you credits on the real data instead, because a fake dataset tells you nothing about whether ours is any good.

Fast: from signup to first API call without talking to us

If it takes six weeks and three calls to get access, it isn’t infrastructure yet.

Fast, for us, is one number: how fast can someone go from “I need creator data” to a working call, without talking to a human? Everything below exists to make that number small.

Monthly plans, with an option for an annual commitment

Start on a monthly plan and cancel any time. Credits roll over, and seats are unlimited on every plan, so putting your whole team on it costs nothing extra. I’d rather earn the renewal every month than lock it into a contract. That said, if you want an annual plan we do those too. But you never have to start there.

Our prices are on the site

Starter is $140 a month billed annually, or $199 month to month. Pro starts at $208. The price per credit is published too, as low as $0.15. You can work out what we cost before you talk to anyone, which is not how most of this category sells data.

An API key on signup

Sign up and your API key is on the page, with ten free credits to test with and the docs at docs.influencers.club. No payment details to start. It’s the same dataset behind the dashboard, the API, our MCP server and custom exports.

Plug it into your stack

Native connections to HubSpot, Salesforce and Snowflake, and nodes for Clay, Zapier and n8n. The n8n node exists because every one of our playbooks used to assume you had engineers to build the integration, and now you don’t need them.

An official ChatGPT connector

influencers.club is an official ChatGPT connector. One click, sign in with your account, and ChatGPT can search and enrich creators like any other tool it has. On claude.ai it’s a hosted connector, and there’s an open-source local server for Cursor, VS Code, Claude Desktop, Windsurf and Zed.

Plug it in and just ask. The answer comes from the actual records, not from a model’s guess, and setup takes about two minutes.

Creatorbooks

There are sixteen of them at last count. Creatorbooks are our way of documenting the best, most clever things companies have actually done using creator data in their real workflows, like finding the creators hiding in your e-commerce customers, screening applications with fake follower checks, or personalizing onboarding when a signup turns out to be a creator. Each one is step by step, with the API calls.

Shared Slack channels

We run shared Slack channels with a lot of the teams building on us, so when something looks off, their engineers talk to our engineers directly instead of going through a ticket.

And the proof it works: Kajabi put creator enrichment into their onboarding flow and lifted onboarding conversions by 10%. Pearpop onboarded 100K+ creators on our data. Both are written up with the implementation, not just the number.

Four ways into the same datasetFIG. 09

Dashboard

Search and filter in the browser, plus natural-language search. Export what you find.

API

Discovery, enrichment and lookup endpoints, used in production onboarding flows and marketplaces.

MCP

Official ChatGPT connector, hosted on claude.ai, open source everywhere else. Ask in plain language.

Exports

Custom datasets and reports, one-off or on a schedule.

HubSpotSalesforceSnowflakeClayZapiern8nMCP
01Sign up. Your API key is on the page.
02Spend the ten free credits on creators you already know.
03Ship the first call. Monthly plan, no annual commitment.
Access and integrations as of September 2026. Source: influencers.club/our-data, influencers.club/pricing

Don’t take my word for any of it

Everything above is checkable. The accuracy numbers live on data.influencers.club with the method next to each one. The sourcing, the edges of the dataset, creator rights and the DPA live on influencers.club/our-data.

When a number changes, the page changes. When we add a data point, it goes on that page with its method.

Every number in this post is published there. Go check.