工具 API

firecrawl_crawl

Starts a crawl job on a website and extracts content from all pages. **Best for:** Extracting content from multiple related pages, when you need comprehensive coverage. **Not recommended for:** Extracting content from a single page (use scrape); when token limits are a concern (use map + batch_scrape); when you need fast results (crawling can be slow). **Warning:** Crawl responses can be very large and may exceed token limits. Limit the crawl depth and number of pages, or use map + batch_scrape for better control. **Common mistakes:** Setting limit or maxDiscoveryDepth too high (causes token overflow) or too low (causes missing pages); using crawl for a single page (use scrape instead). Using a /* wildcard is not recommended. **Prompt Example:** "Get all blog posts from the first two levels of example.com/blog." **Usage Example:** ```json { "name": "firecrawl_crawl", "arguments": { "url": "https://example.com/blog/*", "maxDiscoveryDepth": 5, "limit": 20, "allowExternalLinks": false, "deduplicateSimilarURLs": true, "sitemap": "include" } } ``` **Returns:** Operation ID for status checking; use firecrawl_check_crawl_status to check progress.

其他50 积分

调用信息

工具标识
firecrawl.firecrawl_crawl
服务提供方
Firecrawl
平均响应
0 ms
近 7 天调用
0

输入参数

url必填

Starting URL for the crawl

prompt

Natural language prompt to generate crawler options. Explicitly set parameters will override generated ones.

excludePaths

URL paths to exclude from crawling

includePaths

Only crawl these URL paths

maxDiscoveryDepth

Maximum discovery depth to crawl. The root site and sitemapped pages have depth 0.

sitemap

Sitemap mode when crawling. 'skip' ignores the sitemap entirely, 'include' uses sitemap plus other discovery methods (default), 'only' restricts crawling to sitemap URLs.

limit

Maximum number of pages to crawl (default: 10000)

allowExternalLinks

Allow crawling links to external domains

allowSubdomains

Allow crawling links to subdomains of the main domain

crawlEntireDomain

When true, follow internal links to sibling or parent URLs, not just child paths

delay

Delay in seconds between scrapes to respect site rate limits

maxConcurrency

Maximum number of concurrent scrapes; if unset, team limit is used

webhook

deduplicateSimilarURLs

Remove similar URLs during crawl

ignoreQueryParameters

Do not re-scrape the same path with different (or none) query parameters

scrapeOptions

Options for scraping each page