Skip to content

feat(marketplace): query Bright Data's pre-collected datasets - #30

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace
Open

feat(marketplace): query Bright Data's pre-collected datasets#30
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:feat/marketplace

Conversation

@karaposu

Copy link
Copy Markdown

Adds a marketplace command family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the data pipelines collects.

Why

Every dataset command the CLI has today wraps /datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.

The Python SDK wraps this API in src/brightdata/datasets/; this brings the same capability to the CLI.

The commands

brightdata marketplace list [--featured | --search <text>]
brightdata marketplace fields <dataset>
brightdata marketplace filter --dataset <name> --filter '<json>' --records-limit <n>
brightdata marketplace status <snapshot-id>
brightdata marketplace download <snapshot-id> [--wait]

They map one-to-one onto /datasets/list, /datasets/{id}/metadata, POST /datasets/filter, and /datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.

fields is worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.

$ brightdata marketplace fields linkedin_people_profiles
field          | type   | required | description
about          | text   |          | A concise profile summary...
connections    | number |          | How many connections the profile has
...
46 fields — any of them can be used in a filter.

Two decisions that need your eyes

1. --records-limit is required, though the API treats it as optional.

Omitting it means "no cap" against datasets holding up to 620M records, and cost is only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and an or where and was meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.

2. Datasets are addressed by 48 curated aliases or by raw gd_ id — not by the catalogue's names.

I probed /datasets/list live before designing this. name is a display string, not an identifier:

  • Instagram - Profiles, Facebook - Comments (double space), Manta businesses (trailing space)
  • 1,630 of 1,749 contain spaces; 45 have stray leading/trailing whitespace
  • 43 names are duplicated across idsCrunchbase Ziprecruiter filtered North America maps to 8 different datasets

So name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families; --dataset-id reaches all 1,749; list --search finds an id to paste. Every alias id was validated against the live catalogue (48/48).

If you'd prefer different alias names, they're a single map in one file — easy to change.

Testing

  • 28 new tests (397 total), type-check clean.
  • Verified live against the API, not just mocks: the full filter → poll → download lifecycle, --async hand-off, status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend: cost 0 on both snapshots (used --records-limit 1 on a small dataset).

The live run caught two things mocks couldn't:

  • A zero-match query returns status: failed with warning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.
  • --format jsonl was silently producing pretty-printed JSON: the shared client selects JSON parsing with content_type.includes('application/json'), which is also true for the application/jsonl this endpoint returns. Handled locally here with raw_buffer so output stays byte-exact. Worth noting separately — that substring check in src/utils/client.ts may affect other commands that request jsonl; I didn't change shared behaviour in this PR.

Notes

The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.

Adds a 'marketplace' command family mapping onto those five endpoints:

- list [--featured|--search]  discover datasets (the catalogue is ~1,750
  rows, so a bare list is filtered rather than dumped)
- fields <dataset>           per-field types and descriptions from the
  metadata endpoint — the CLI's first self-documenting data command; you
  can see what is filterable before spending anything
- filter                     the query itself; the filter tree is sent
  verbatim and the API validates it
- status <snapshot-id>       status plus records, file size and cost
- download <snapshot-id>     [--wait]; exits 3 when not ready yet, so
  scripts can tell 'retry later' from 'failed'

Notes on two decisions:

--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.

Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook -  Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.

Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).

28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant