Bright Data
Search, Crawl and Scrape any site, at scale, without getting blocked
1.0.2Bright Data is a web data infrastructure provider; this toolkit enables Arcade agents to scrape pages, run search-engine queries, and extract structured data from major platforms at scale without getting blocked.
Capabilities
- Web scraping: Fetch any public webpage and receive clean Markdown output, suitable for LLM consumption or downstream processing.
- Search engines: Query Google, Bing, or Yandex with configurable parameters — result count, country code, search type (web, images, etc.).
- Structured data feeds: Pull pre-parsed, schema'd records from platforms including Amazon, LinkedIn, Instagram, Facebook, X, YouTube, Zillow, Booking.com, and ZoomInfo — no custom parser needed.
Secrets
No OAuth flow is used. Access is authenticated via two required secrets passed to the toolkit.
BRIGHTDATA_API_KEY
Your Bright Data account API key. Obtain it from the Bright Data control panel: log in, navigate to Account Settings → API Tokens, and generate or copy your token. The key authenticates all API requests made on your behalf.
BRIGHTDATA_ZONE
The Bright Data zone (proxy zone or dataset zone) to route requests through. Zones are created and named in the Bright Data control panel under Proxies & Scraping Infrastructure. Use the exact zone name as it appears in your dashboard (e.g., residential, datacenter, or a custom zone name you have configured). The zone determines the proxy pool, geolocation behavior, and billing ruleset applied to each request.
For help configuring secrets in Arcade, see the Arcade secrets guide. You can also manage secrets at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(3)
| Tool name | Description | Secrets | |
|---|---|---|---|
Scrape a webpage and return content in Markdown format using Bright Data.
Examples:
scrape_as_markdown("https://example.com") -> "# Example Page
Content..."
scrape_as_markdown("https://news.ycombinator.com") -> "# Hacker News
..."
| 2 | ||
Search using Google, Bing, or Yandex with advanced parameters using Bright Data.
Examples:
search_engine("climate change") -> "# Search Results
## Climate Change - Wikipedia
..."
search_engine("Python tutorials", engine="bing", num_results=5) -> "# Bing Results
..."
search_engine("cats", search_type="images", country_code="us") -> "# Image Results
..."
| 2 | ||
Extract structured data from various websites like LinkedIn, Amazon, Instagram, etc.
NEVER MAKE UP LINKS. IF LINKS ARE NEEDED, FIND THEM WITH A WEB SEARCH FIRST.
Supported source types:
- amazon_product, amazon_product_reviews
- linkedin_person_profile, linkedin_company_profile
- zoominfo_company_profile
- instagram_profiles, instagram_posts, instagram_reels, instagram_comments
- facebook_posts, facebook_marketplace_listings, facebook_company_reviews
- x_posts
- zillow_properties_listing
- booking_hotel_listings
- youtube_videos
Examples:
web_data_feed("amazon_product", "https://amazon.com/dp/B08N5WRWNW")
-> "{"title": "Product Name", ...}"
web_data_feed("linkedin_person_profile", "https://linkedin.com/in/johndoe")
-> "{"name": "John Doe", ...}"
web_data_feed(
"facebook_company_reviews", "https://facebook.com/company", num_of_reviews=50
) -> "[{"review": "...", ...}]" | 1 |