Get paid for fan interactions — start free.

Create your free FanBell link

Wishlist / Project Support

Fund an Open Dataset Project From Fans

How researchers, developers, and data journalists turn 'someone should build this' into a funded open dataset release — a Wishlist / Project Support goal with a progress bar, not a paid review of someone else's data.

Updated August 2026

Get paid for this — with FanBell

Absorbing the collection, cleaning, and licensing unpaid? Get funded by the people who'll use the dataset.

FanBell is a link in your bio where fans pay you directly for:

Wishlist62%Custom service$120Paid question$25Shoutout$60Tip$5+

Post the release as a Project Support goal and a progress bar fills as backers cover the collection and cleaning — 12% only when it does.

No monthly fee · 12% only when a fan pays

Fund an open dataset project by setting a Wishlist / Project Support goal for one specific dataset — its scope, collection method, and release license — paired with a visible progress bar, so the people who'd use the finished dataset can back the collection and cleaning work directly instead of you absorbing it unpaid.

FanBell's pricing page states a "12% platform fee per paid transaction · $0/month" (fanbell.link/pricing).

Public datasets are rarely built by the people who need them most. One person usually absorbs the collection, cleaning, and documentation cost, and everyone else downloads the finished file for free. The Hugging Face Hub documentation states the Hub "is home to over 500k public datasets in more than 8k languages" (Hugging Face Hub documentation). Common Crawl reports that its archive totals more than 10 petabytes, has been cited in over 12,000 research papers, and that 64% of large language models are trained on Common Crawl data. Free at the point of download says nothing about who paid to build it — a funding goal moves that cost to the people who will actually use the result.

What makes an open dataset project fundable?

An open dataset project is fundable when it names one bounded dataset: what it covers, how it is collected, what format it ships in, which license it carries, and how much money finishes it. Backers fund a file they can picture, so the scope has to state a point at which the goal is met and the data is published.

"A labeled set of 5,000 satellite images of wildfire smoke plumes" or "a cleaned CSV of state minimum-wage history since 1970" are fundable because each has a defined end state a backer can check against. "Build up a general data collection over time" is not, since there is no point at which the goal is met.

Scope should also state what "released" means: a public download link, a stated license, and documentation of the collection method. A dataset that is funded but never published is not a finished project — it is a private file someone else paid for.

What's the difference between funding a dataset and accepting tips?

A tip is an unscoped one-time payment with no stated purpose beyond thanks. A dataset-funding goal is scoped to one named release, states a target amount, and shows a progress bar toward that target. Both offers can run side by side on the same creator page, and neither one requires the other to work.

On FanBell, Tips cover open-ended one-time support with no reply or delivery required — the same mechanic an open source maintainer collecting tips would use for general thanks — while Wishlist / Project Support is cash directed at a named goal. FanBell's own feature page describes Project Support as "direct cash toward something you're working on. You describe the goal, set a target amount, and fans contribute — with a progress bar that fills as support comes in". Someone dropping $5 because they like your data-journalism newsletter is a tip; forty people putting money toward "$2,500 to clean and release the county permitting dataset" is a project goal.

Where do you host and license the dataset once it's built?

Hosting is where the finished file lives and how other people cite it. Licensing is the separate legal question of what downstream users may do with that file. Zenodo, the Hugging Face Hub, and GitHub all accept public dataset uploads at no cost, each with published size limits. Data.gov is different: it accepts no uploads from individuals at all.

HostBest forPersistent citationDocumented limits and cost
ZenodoResearch datasets that need a citable DOIRegisters a DOI for every published uploadFree; 50 GB total files and a maximum of 100 files per record (Zenodo Policies)
Hugging Face HubML/NLP datasets loaded directly into training codeRepo URL; a DOI can be generated from repo settings, not automaticallyFree public storage on a best-effort basis; no per-repo size limit for datasets, with a 500 GB hard limit on any single file (Hugging Face — Storage limits)
GitHubSmaller files released alongside codeRepo URL; pairs with a Zenodo DOI through Zenodo's GitHub integrationUnlimited public repositories on GitHub Free; files over 100 MiB are blocked (GitHub Docs — About large files)
Data.govFederal, state, local, and tribal agency catalogs — not individual researchersCatalog listing, not a DOINo individual submission path; an agency publishes a metadata file that Data.gov harvests on a schedule (resources.data.gov)

The Data.gov row is the one most people get wrong, so it is worth stating plainly what the government's own publisher guidance says.

"Data.gov is the federal government's central open data catalog. It does not host data files directly. Instead, it collects metadata (the descriptions, contact information, and download links for your datasets) from a file your agency publishes and maintains." — resources.data.gov, How to get your Open Data on Data.gov

In practice that means an independent dataset builder lists on Zenodo, Hugging Face, or GitHub, and only reaches Data.gov if a government agency adopts the data and adds it to a harvest source. Data.gov's catalog carried 556,474 datasets at the time of checking, all of them harvested from agency-published catalogs rather than uploaded by members of the public.

GitHub Docs give the two hard numbers that decide whether a dataset can live in a repo at all: "GitHub blocks files larger than 100 MiB" and "We recommend repositories remain small, ideally less than 1 GB, and less than 5 GB is strongly recommended". Anything past that belongs on Zenodo or the Hugging Face Hub.

Zenodo was launched in May 2013 by CERN with OpenAIRE as a free, general-purpose repository for citable research outputs, including datasets. Zenodo also registers a DOI for every published record and issues a separate version-specific DOI each time you publish a new version, so a corrected dataset stays citable without breaking the old citation (Zenodo — DOI versioning). Common Crawl shows what a sustained open-data release looks like at the far end of that scale.

"Common Crawl is a nonprofit 501(c)(3) organisation that crawls the web and freely provides its archives and datasets to the public." — Common Crawl, About

Which license should an open dataset use?

For an independently funded dataset the practical choice is usually between two well-documented options: Creative Commons CC0, which waives copyright and related rights entirely, and the Open Data Commons Open Database License (ODbL), a copyleft license that requires attribution and keeps redistributed copies under the same terms. Pick one before funding opens.

Creative Commons describes CC0 as "a universal legal tool that allows creators and rightsholders to waive all copyright and related rights in their works to the fullest extent permitted by law" (Creative Commons — Public Domain). Open Data Commons describes the ODbL as a license "intended to allow users to freely share, modify, and use this Database while maintaining this same freedom for others". The largest well-known example of the copyleft route is mapping data: the OpenStreetMap Foundation states that data extracted from OpenStreetMap after September 2012 is licensed under the Open Database License 1.0 (OpenStreetMap Foundation — Licence).

Pick the license before funding opens. Backers are effectively funding the terms under which the result gets used, and changing the license after release is far harder than choosing correctly the first time.

Do you need a data management plan or DOI before funding starts?

Only when a specific funder or institution requires one. A formal data management plan is a grant-application requirement, not a general condition of publishing open data, so an independently funded dataset release does not create one on its own. A DOI is likewise optional — useful for academic reuse, not required for publication.

The clearest example of a funder-imposed requirement is the U.S. National Science Foundation, which uses the term Data Management and Sharing Plan (DMSP) rather than "data management plan." NSF states: "The two-page data management and sharing plan is a required part of a proposal to the U.S. National Science Foundation". The NSF Proposal and Award Policies and Procedures Guide (PAPPG 24-1), Chapter II, sets the same two-page cap on the Data Management and Sharing Plan uploaded as a supplementary document. That requirement attaches to NSF proposals, not to crowd-funded projects — but the questions it asks (what is collected, how it is stored, under what license) are worth answering on any funding page.

A DOI from a repository like Zenodo at release is a separate, optional step that makes a dataset easier to cite. It is not required, but it is worth naming on the goal page if the dataset targets academic reuse.

How does a funding goal with a progress bar work on FanBell?

A FanBell Wishlist / Project Support goal is a page where fans send money toward a stated target, with a progress bar that fills as contributions come in. It is cash toward a goal, not a store: no shipping, no tiered reward levels, and no product being fulfilled. Setup is free and there is no follower minimum.

FanBell's own documentation is the authority on its mechanics. The pricing page states that "the fan pays only the displayed price. FanBell charges a 12% platform fee, and payment-processing fees are deducted separately from creator earnings", and FanBell's machine-readable summary states plainly: "No follower minimum — any creator can set up a page". A connected Stripe account is required before any paid widget can take money (fanbell.link/how-it-works).

Setting one up means naming the dataset, writing what it covers and what "released" looks like, and linking the page from wherever people already find you — a repo README, a newsletter, or a pinned post. For what fields a strong goal page includes, see what to put on a project support page.

How do you price and scope a dataset-funding goal?

Price a dataset goal against the real cost of the work it covers — collection time, cleaning, labeling, documentation, and hosting — rather than picking a round number with no stated basis. A goal a backer can map to concrete tasks reads as more credible than an unexplained total, and it answers the obvious question of what the money buys.

Task coveredExample goal framingWhat it funds
Manual labeling"$1,200 to label 5,000 images by hand"Labeling time on a bounded image set
Cleaning + documentation"$600 to clean and document the raw export"Turning a messy raw file into a documented release
Hosting + DOI + write-up"$500 for a stable DOI and a release write-up"Zenodo deposit costs $0 under its published policies, but writing docs and metadata takes real time
Transcription"$900 to transcribe 40 hours of archival audio"Converting raw recordings into a searchable text dataset

These figures are illustrative examples of how to frame a target, not earnings guarantees — actual amounts depend on the dataset and how clearly the scope is stated. Budget for the payment fees too: Stripe's published US pricing is 2.9% + $0.30 per successful transaction for domestic cards (Stripe pricing), charged on top of FanBell's 12% platform fee. On a $2,500 goal that is roughly $300 in platform fee plus $80 in Stripe fees across 25 contributions of $100 (2.9% of $2,500 = $72.50, plus $0.30 × 25 = $7.50). Funding a bug bounty or CVE research project walks through the same milestone-framing approach for a security-research goal.

Where does dataset funding fit next to a paid dataset review?

Funding a dataset and getting paid to review one are two different offers. A Wishlist / Project Support goal funds building and releasing a new dataset of your own. A paid dataset review is a priced Creator Service where a fan sends you a dataset they already have and you return feedback on methodology and conclusions.

A researcher or developer can run both without conflict — a funding goal for a dataset you are building, and a paid review service for feedback on data others have already collected. See selling a dataset or analysis review for the review side; for how async paid work fits alongside a technical creator's other income, see FanBell for developers; for the same funding-goal mechanic applied to an open source roadmap, see funding your open source project's next milestone; the same milestone-based scoping also works for funding a UX case study project outside the data world.

Frequently asked questions

Is dataset funding the same as accepting tips?

No. A tip is an unscoped one-time payment with no stated purpose, while a Wishlist / Project Support goal names a specific dataset, states a target amount, and shows "a progress bar that fills as support comes in". Both can run on the same FanBell page as separate offers.

Can I list my funded dataset on Data.gov?

Not directly. Data.gov "does not host data files directly" and instead harvests metadata from a catalog file that a government agency publishes and maintains (resources.data.gov). An independently funded dataset goes on Zenodo, the Hugging Face Hub, or GitHub, and reaches Data.gov only if an agency adopts it into a harvest source.

Do I need a nonprofit or research institution to run a dataset-funding goal?

No. FanBell requires a connected Stripe account for paid widgets but no incorporation or institutional affiliation, and sets no follower minimum. A data management plan requirement, where it applies, comes from a specific funder such as the NSF, whose two-page Data Management and Sharing Plan is required only for NSF proposals.

Which license should I pick if I'm not sure?

CC0 is the simpler default for maximum reuse, described by Creative Commons as a tool to "waive all copyright and related rights in their works to the fullest extent permitted by law". The Open Data Commons ODbL 1.0 is the copyleft option when you want redistributed copies to stay open. Decide before funding opens, since backers are also funding the terms of reuse.

How big can the dataset file actually be?

It depends on the host, and each publishes a limit. Zenodo accepts 50 GB total and a maximum of 100 files per record. GitHub blocks any file larger than 100 MiB. The Hugging Face Hub sets no per-repo size limit for datasets but caps a single file at 500 GB.

What does FanBell charge to run a dataset-funding goal?

FanBell's pricing page lists a "12% platform fee per paid transaction · $0/month," with payment-processing fees deducted separately from creator earnings. Stripe's published US rate for domestic cards is 2.9% + $0.30 per successful transaction.

Create your free FanBell page and let the people who'd use your dataset fund the work of actually building it.

Ready to get paid for the interactions you already get?

Create your free FanBell link