feat(marketplace): query Bright Data's pre-collected datasets - #30
Open
karaposu wants to merge 1 commit into
Open
feat(marketplace): query Bright Data's pre-collected datasets#30karaposu wants to merge 1 commit into
karaposu wants to merge 1 commit into
Conversation
The CLI could only ever collect NEW data: pipelines wraps /datasets/v3/*,
which scrapes what you point it at and bills accordingly. Bright Data also
sells ~1,750 pre-collected datasets queried through a separate, unversioned
API (/datasets/list, /datasets/{id}/metadata, POST /datasets/filter,
/datasets/snapshots/{id}[/download]) that the CLI did not touch at all —
so users had to re-scrape data Bright Data already held, and had no way to
discover that these datasets existed.
Adds a 'marketplace' command family mapping onto those five endpoints:
- list [--featured|--search] discover datasets (the catalogue is ~1,750
rows, so a bare list is filtered rather than dumped)
- fields <dataset> per-field types and descriptions from the
metadata endpoint — the CLI's first self-documenting data command; you
can see what is filterable before spending anything
- filter the query itself; the filter tree is sent
verbatim and the API validates it
- status <snapshot-id> status plus records, file size and cost
- download <snapshot-id> [--wait]; exits 3 when not ready yet, so
scripts can tell 'retry later' from 'failed'
Notes on two decisions:
--records-limit is REQUIRED, though the API treats it as optional. Omitting
it means no cap on datasets holding hundreds of millions of records, and
cost is only observable after the query is committed, so there is no way to
preview or undo an accidentally huge query. Making it explicit costs one
flag and removes the whole unbounded case.
Datasets are addressed by 48 curated aliases or by raw gd_ id, not by the
catalogue's own names. A live probe showed why: 'name' is a display string
('Instagram - Profiles', 'Facebook - Comments' with a double space,
'Manta businesses ' with a trailing one), 1630 of 1749 contain spaces, and
43 names are shared across different ids — one name maps to 8 datasets. So
names are for reading; --search finds an id to paste. All 48 alias ids were
validated against the live catalogue.
Also adds src/utils/exit-codes.ts (0 ok / 1 error / 3 not ready).
28 new tests (397 total). Verified live against the API: the full
filter -> poll -> download lifecycle, all three formats, and every guard
firing before any network call. All 48 alias ids validated against the
live catalogue.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a
marketplacecommand family, giving the CLI access to Bright Data's Dataset Marketplace — the ~1,750 pre-collected datasets you query, as opposed to the datapipelinescollects.Why
Every dataset command the CLI has today wraps
/datasets/v3/*: you supply inputs, Bright Data scrapes them, you're billed for scraping. The marketplace is a different product on a different (unversioned) API family that the CLI didn't touch at all — so a user wanting, say, 1,000 Technology-industry company records had to scrape them, even though Bright Data already holds them and will serve them from a filter query. There was also no way to discover from the CLI that these datasets exist.The Python SDK wraps this API in
src/brightdata/datasets/; this brings the same capability to the CLI.The commands
They map one-to-one onto
/datasets/list,/datasets/{id}/metadata,POST /datasets/filter, and/datasets/snapshots/{id}[/download], and reuse the existing client, poller, format and output plumbing.fieldsis worth calling out: it returns each field's type and description straight from the metadata endpoint, so you can see what's filterable before spending anything. No other CLI command is self-documenting in that way.Two decisions that need your eyes
1.
--records-limitis required, though the API treats it as optional.Omitting it means "no cap" against datasets holding up to 620M records, and
costis only observable on the snapshot after the query has been committed — there's no estimate endpoint, so the CLI cannot show a price first. A filter tree is nested JSON typed on a command line, and anorwhereandwas meant widens the match by orders of magnitude. Making the cap explicit costs one flag and removes the entire unbounded case. Happy to relax it if you'd rather.2. Datasets are addressed by 48 curated aliases or by raw
gd_id — not by the catalogue's names.I probed
/datasets/listlive before designing this.nameis a display string, not an identifier:Instagram - Profiles,Facebook - Comments(double space),Manta businesses(trailing space)Crunchbase Ziprecruiter filtered North Americamaps to 8 different datasetsSo name-based resolution can't identify a dataset even on an exact match. The 48 aliases cover the social/content and business-intelligence families;
--dataset-idreaches all 1,749;list --searchfinds an id to paste. Every alias id was validated against the live catalogue (48/48).If you'd prefer different alias names, they're a single map in one file — easy to change.
Testing
filter→ poll →downloadlifecycle,--asynchand-off,status, all three formats (json/csv/jsonl), and every guard firing before any network call. Total spend:cost 0on both snapshots (used--records-limit 1on a small dataset).The live run caught two things mocks couldn't:
status: failedwithwarning_code: no_records_found, not a ready-but-empty snapshot. It's now reported as "The filter matched no records" with exit 0, rather than as a failure.--format jsonlwas silently producing pretty-printed JSON: the shared client selects JSON parsing withcontent_type.includes('application/json'), which is also true for theapplication/jsonlthis endpoint returns. Handled locally here withraw_bufferso output stays byte-exact. Worth noting separately — that substring check insrc/utils/client.tsmay affect other commands that request jsonl; I didn't change shared behaviour in this PR.Notes
index.ts. No existing behaviour touched.src/utils/exit-codes.ts(0ok /1error /3not ready) —downloadexits3when a snapshot is still building, so scripts can tell "retry later" from "failed".