Bright Data
Search, Crawl and Scrape any site, at scale, without getting blocked
0.5.1Bright Data is a web data platform that provides infrastructure for searching, crawling, and scraping the web at scale without being blocked. The Arcade toolkit exposes Bright Data's capabilities as callable tools for web content retrieval, search, and structured data extraction.
Capabilities
- Web scraping: Fetch any public webpage and receive the content normalized as Markdown, suitable for downstream LLM processing or analysis.
- Multi-engine search: Run queries against Google, Bing, or Yandex with configurable result count, search type (web, images), and country targeting.
- Structured data feeds: Extract pre-structured JSON records from major platforms — including Amazon (products, reviews), LinkedIn (people, companies), Instagram (profiles, posts, reels, comments), Facebook (posts, marketplace, reviews), X/Twitter, Zillow, Booking.com, YouTube, and ZoomInfo — without writing custom scrapers.
Secrets
This toolkit requires two secrets to authenticate with Bright Data's API.
-
BRIGHTDATA_API_KEY— Your Bright Data API token, used to authenticate all requests. Obtain it from the Bright Data control panel under Account Settings → API Token. A paid or trial Bright Data account is required; the key is account-scoped and grants access to all zones and datasets associated with your account. -
BRIGHTDATA_ZONE— The name of the Bright Data proxy zone or dataset zone to route requests through (e.g.,"residential","datacenter", or a custom zone name). Zones are created and managed in the Bright Data control panel under Proxies & Scraping Infrastructure. The zone determines the proxy pool, geolocation options, and rate limits available to your requests. Use the exact zone name string shown in your dashboard.
For guidance on configuring secrets in Arcade, see the Arcade secrets documentation. You can also manage secrets directly at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(3)
| Tool name | Description | Secrets | |
|---|---|---|---|
Scrape a webpage and return content in Markdown format using Bright Data.
Examples:
scrape_as_markdown("https://example.com") -> "# Example Page
Content..."
scrape_as_markdown("https://news.ycombinator.com") -> "# Hacker News
..."
| 2 | ||
Search using Google, Bing, or Yandex with advanced parameters using Bright Data.
Examples:
search_engine("climate change") -> "# Search Results
## Climate Change - Wikipedia
..."
search_engine("Python tutorials", engine="bing", num_results=5) -> "# Bing Results
..."
search_engine("cats", search_type="images", country_code="us") -> "# Image Results
..."
| 2 | ||
Extract structured data from various websites like LinkedIn, Amazon, Instagram, etc.
NEVER MADE UP LINKS - IF LINKS ARE NEEDED, EXECUTE search_engine FIRST.
Supported source types:
- amazon_product, amazon_product_reviews
- linkedin_person_profile, linkedin_company_profile
- zoominfo_company_profile
- instagram_profiles, instagram_posts, instagram_reels, instagram_comments
- facebook_posts, facebook_marketplace_listings, facebook_company_reviews
- x_posts
- zillow_properties_listing
- booking_hotel_listings
- youtube_videos
Examples:
web_data_feed("amazon_product", "https://amazon.com/dp/B08N5WRWNW")
-> "{"title": "Product Name", ...}"
web_data_feed("linkedin_person_profile", "https://linkedin.com/in/johndoe")
-> "{"name": "John Doe", ...}"
web_data_feed(
"facebook_company_reviews", "https://facebook.com/company", num_of_reviews=50
) -> "[{"review": "...", ...}]" | 2 |