Firecrawl is a powerful web scraping and content extraction API that integrates seamlessly into Sim, enabling developers to extract clean, structured content from any website. This integration provides a simple way to transform web pages into usable data formats like Markdown and HTML while preserving the essential content.
With Firecrawl in Sim, you can:
- Extract clean content: Remove ads, navigation elements, and other distractions to get just the main content
- Convert to structured formats: Transform web pages into Markdown, HTML, or JSON
- Capture metadata: Extract SEO metadata, Open Graph tags, and other page information
- Handle JavaScript-heavy sites: Process content from modern web applications that rely on JavaScript
- Filter content: Focus on specific parts of a page using CSS selectors
- Process at scale: Handle high-volume scraping needs with a reliable API
- Search the web: Perform intelligent web searches and retrieve structured results
- Crawl entire sites: Crawl multiple pages from a website and aggregate their content
In Sim, the Firecrawl integration enables your agents to access and process web content programmatically as part of their workflows. Supported operations include:
- Scrape: Extract structured content (Markdown, HTML, metadata) from a single web page.
- Search: Search the web for information using Firecrawl's intelligent search capabilities.
- Crawl: Crawl multiple pages from a website, returning structured content and metadata for each page.
This allows your agents to gather information from websites, extract structured data, and use that information to make decisions or generate insights—all without having to navigate the complexities of raw HTML parsing or browser automation. Simply configure the Firecrawl block with your API key, select the operation (Scrape, Search, or Crawl), and provide the relevant parameters. Your agents can immediately begin working with web content in a clean, structured format.
Usage Instructions
Integrate Firecrawl into the workflow. Scrape pages, search the web, crawl entire sites, map URL structures, and extract structured data with AI.
Tools
firecrawl_scrape
Extract structured content from web pages with comprehensive metadata support. Converts content to markdown or HTML while capturing SEO metadata, Open Graph tags, and page information.
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
url | string | Yes | The URL to scrape content from |
scrapeOptions | json | No | Options for content scraping |
apiKey | string | Yes | Firecrawl API key |
Output
| Parameter | Type | Description |
|---|---|---|
markdown | string | Page content in markdown format |
html | string | Raw HTML content of the page |
metadata | object | Page metadata including SEO and Open Graph information |
firecrawl_search
Search for information on the web using Firecrawl
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
query | string | Yes | The search query to use |
apiKey | string | Yes | Firecrawl API key |
Output
| Parameter | Type | Description |
|---|---|---|
data | array | Search results data |
firecrawl_crawl
Crawl entire websites and extract structured content from all accessible pages
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
url | string | Yes | The website URL to crawl |
limit | number | No | Maximum number of pages to crawl (default: 100) |
onlyMainContent | boolean | No | Extract only main content from pages |
apiKey | string | Yes | Firecrawl API Key |
Output
| Parameter | Type | Description |
|---|---|---|
pages | array | Array of crawled pages with their content and metadata |
firecrawl_map
Get a complete list of URLs from any website quickly and reliably. Useful for discovering all pages on a site without crawling them.
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
url | string | Yes | The base URL to map and discover links from |
search | string | No | Filter results by relevance to a search term (e.g., "blog") |
sitemap | string | No | Controls sitemap usage: "skip", "include" (default), or "only" |
includeSubdomains | boolean | No | Whether to include URLs from subdomains (default: true) |
ignoreQueryParameters | boolean | No | Exclude URLs containing query strings (default: true) |
limit | number | No | Maximum number of links to return (max: 100,000, default: 5,000) |
timeout | number | No | Request timeout in milliseconds |
location | json | No | Geographic context for proxying (country, languages) |
apiKey | string | Yes | Firecrawl API key |
Output
| Parameter | Type | Description |
|---|---|---|
success | boolean | Whether the mapping operation was successful |
links | array | Array of discovered URLs from the website |
firecrawl_extract
Extract structured data from entire webpages using natural language prompts and JSON schema. Powerful agentic feature for intelligent data extraction.
Input
| Parameter | Type | Required | Description |
|---|---|---|---|
urls | json | Yes | Array of URLs to extract data from (supports glob format) |
prompt | string | No | Natural language guidance for the extraction process |
schema | json | No | JSON Schema defining the structure of data to extract |
enableWebSearch | boolean | No | Enable web search to find supplementary information (default: false) |
ignoreSitemap | boolean | No | Ignore sitemap.xml files during scanning (default: false) |
includeSubdomains | boolean | No | Extend scanning to subdomains (default: true) |
showSources | boolean | No | Return data sources in the response (default: false) |
ignoreInvalidURLs | boolean | No | Skip invalid URLs in the array (default: true) |
scrapeOptions | json | No | Advanced scraping configuration options |
apiKey | string | Yes | Firecrawl API key |
Output
| Parameter | Type | Description |
|---|---|---|
success | boolean | Whether the extraction operation was successful |
data | object | Extracted structured data according to the schema or prompt |
Notes
- Category:
tools - Type:
firecrawl