Skip to content
Model Context Protocol · v0.3.1

Web data. Browser actions.
One MCP connection.

Let your agent read pages, extract structured data, interact with browser-rendered content and process up to 100 URLs in a batch. Nine tools cover the workflow — including authorized HTTP writes, background jobs and token usage. Markdown by default; only successful scrapes are billed.

Run it locally over stdio with npx or pip, or connect to the hosted endpoint over Streamable HTTP and sign in with OAuth instead of pasting a key.

Choose how to connect

Use the hosted endpoint if your client supports remote MCP. Prefer a local process? Use the npm or Python package with your API key. You need a FineData account either way.

Hosted endpoint · Streamable HTTP

https://mcp.finedata.ai/mcp

In clients with a remote connector UI, add this URL and choose OAuth sign-in. The JSON examples below use Cursor's mcpServers shape. Other clients may use a connector UI or CLI instead; there is no universal config file.

Local, with a key

The agent runs the server as a child process. Best for an agent on your own machine.

{
  "mcpServers": {
    "finedata": {
      "command": "npx",
      "args": ["-y", "@finedata/mcp-server"],
      "env": { "FINEDATA_API_KEY": "fd_your_api_key" }
    }
  }
}

Requires Node.js and npm. In Cursor, use ~/.cursor/mcp.json; Claude Desktop accepts this local entry in claude_desktop_config.json. Merge it with existing servers. For Python, run pip install finedata-mcp, then set command to finedata-mcp and args to [], keeping the same env.

Remote, with a key

Nothing to install or update. Best for hosted agents and CI.

{
  "mcpServers": {
    "finedata": {
      "url": "https://mcp.finedata.ai/mcp",
      "headers": { "Authorization": "Bearer fd_your_api_key" }
    }
  }
}

Keys come from your dashboard and carry your own plan and balance.

Remote, with sign-in

No credentials in the config at all. Best when someone else runs the agent.

{
  "mcpServers": {
    "finedata": {
      "url": "https://mcp.finedata.ai/mcp"
    }
  }
}

The client registers itself and opens a browser to sign in. See how that works.

Make one call before building a loop

Reload your client's MCP connection and check that nine tools appear. Ask your agent:

Use FineData to read https://example.com as markdown. Report the page title and the tokens used. Do not enable extra rendering modes.

Tool: scrape_url

{
  "url": "https://example.com",
  "formats": ["markdown"],
  "only_main_content": true
}

What comes back

A text result with the URL, target status, tokens used and page content. Structured extraction is appended as JSON; screenshots are returned as MCP image content. This is not the REST API's JSON envelope.

The input is a URL, not a search phrase. To read a site's search results, pass its search URL with an encoded query string. There is no query argument or general web-search tool.

More than page-to-markdown

Use the same connection for structured extraction, multi-page research and browser interactions. These are tool arguments, not client configuration.

Extract the fields you need

Turn a product page into structured data rather than asking the agent to parse a wall of HTML. Replace this illustrative URL with a real product page.

Tool: scrape_url

{
  "url": "https://example.com/product",
  "extract_schema": {
    "type": "object",
    "properties": {
      "name": {
        "type": "string"
      },
      "price": {
        "type": "number"
      },
      "currency": {
        "type": "string"
      }
    },
    "required": [
      "name",
      "price",
      "currency"
    ]
  }
}
Read many pages in one job

Submit up to 100 URLs and save both batch_id and job_ids. Track progress with get_batch_status, then fetch each result with get_job_status. Per-URL objects can override shared options.

Tool: batch_scrape

{
  "urls": [
    "https://example.com",
    "https://example.org"
  ],
  "formats": [
    "markdown"
  ],
  "only_main_content": true
}
Wait for browser-rendered content

Render JavaScript and wait for a specific element. js_actions can also click, type, scroll and evaluate page JavaScript. Replace the illustrative URL and selector.

Tool: scrape_url

{
  "url": "https://example.com/catalog",
  "use_js_render": true,
  "js_wait_for": "selector:.product",
  "formats": [
    "markdown"
  ]
}
Send an authorized write request

Use send_http_request for POST, PUT, PATCH or DELETE. This can change data on the target site. Obtain permission first; the following URL is a placeholder for an endpoint you control. Do not blindly retry non-idempotent writes.

{
  "url": "https://your-api.example/items",
  "method": "POST",
  "headers": {
    "Content-Type": "application/json"
  },
  "body": "{\"name\":\"Example item\"}",
  "auto_retry": false,
  "formats": [
    "text"
  ]
}

For slow pages, use scrape_async, save its job_id, then poll get_job_status with a delay. Stop on completed, failed or cancelled. MCP async and batch tools are GET-only; writes use the synchronous send_http_request tool.

The nine tools

Tool annotations describe read-only, destructive and idempotent behavior; they are hints for the client, not permission checks or a budget cap. scrape_url sends GET; send_http_request handles HTTP writes. Browser actions such as clicking or typing can also change a site, even during a GET-based scrape. Review those actions and your client's confirmation settings.

Tool What it does Scope
scrape_url Read a URL with GET. Return markdown, extraction data or screenshots; supports rendering and browser actions. scrape:write
send_http_request Send POST, PUT, PATCH or DELETE through the same pipeline, for forms and write APIs. scrape:write
scrape_async Queue one GET request and receive a job_id. Poll with get_job_status. scrape:write
get_job_status Poll one queued job and read its result when finished. jobs:read
list_jobs List recent jobs with their state, to resume after a restart. jobs:read
cancel_job Cancel a pending job. Processing jobs cannot be cancelled. scrape:write
batch_scrape Submit 1–100 URLs with shared options and per-URL overrides. GET-only. scrape:write
get_batch_status Track overall batch progress. Fetch individual results with get_job_status using saved job_ids. jobs:read
get_usage Read token consumption for the current period. Not a balance, plan limit or request log. usage:read

Use list_jobs to resume asynchronous work, not to retrieve synchronous scrape history. Only pending jobs can be cancelled; processing jobs finish normally. For complete field definitions, see the API options and your client's tool schema.

Sign-in instead of a shared key

An API key is a bearer secret: whoever holds it is you. That is fine for an agent you run yourself and awkward for one you hand to a teammate or a customer. The remote endpoint therefore speaks OAuth 2.1, so each person connects their own account.

What happens

  1. The client discovers the endpoint's metadata and registers itself — no client id to create by hand.
  2. A browser opens on a consent screen naming the client and the scopes it asks for.
  3. You sign in to your own account and approve. The client receives a short-lived token, verified with PKCE so an intercepted redirect is useless.
  4. Requests bill to your account. The token lasts an hour and renews quietly for up to 30 days.

What the token cannot do

Approving "scrape on my behalf" grants exactly that. The token reaches the nine tools and their permitted API endpoints: it cannot create API keys, change your plan, manage wallet operations or read your profile. Scrapes still consume tokens and can incur normal usage charges on your account.

The hosted MCP connection requests all three scopes listed above. The API enforces them per endpoint. Revoke a connection from your dashboard; changing your password also invalidates existing connections.

Why cost matters more for agents

Failures are not billed

An agent retrying on its own judgement would otherwise be an open budget. You pay for requests that returned the page.

The agent can see the bill

Successful scrape results report tokens used. get_usage reports current-period consumption, not your remaining balance or plan limit. Check limits in the billing dashboard and set a separate budget for your agent.

Markdown by default

Raw HTML burns context for nothing. Long pages are truncated with a note rather than silently cut, so the model knows what it did not see.

A default fetch is 3 tokens before retries, extraction or rendering. Start there, add use_js_render for JavaScript content, and use failure hints to decide whether another rendering mode is appropriate. No mode guarantees access to a particular site.

Text content is truncated at about 60,000 characters with a note. Use only_main_content or structured extraction for smaller results. For CSV/XLSX exports or content-only responses, use the REST API; MCP is designed for agent-readable text and images.

If the first call fails

No tools appear
Check the client's MCP logs. For local setup, verify Node.js/npm is available to the app and the key is in env. Reload the connection. For remote setup, use the full URL ending in /mcp and a client with Streamable HTTP support.
401 or OAuth sign-in failure
Verify the key or reconnect with OAuth. A remote API key belongs in Authorization: Bearer fd_…, not in the URL. Keep keys out of shared repositories and screenshots.
402, 403 or 429
Read the error: check billing for 402, account or feature access for 403, and reduce request frequency for 429. Changing the rendering engine does not fix these errors.
A job is still running
Keep the returned job_id or batch_id and poll that job with a delay. Do not submit the same work again. A batch can finish with partial failures; inspect individual results.

Questions

What is the FineData MCP server?

It is a Model Context Protocol server with nine tools for reading pages, extracting structured data, rendering JavaScript, running browser actions, sending HTTP writes, managing asynchronous jobs and batches, and checking token consumption. It runs locally over stdio via npx or pip, or remotely over Streamable HTTP at https://mcp.finedata.ai/mcp. Pages come back as markdown by default, which is what a language model reads best.

How do I connect FineData MCP to Cursor or Claude Desktop?

For local stdio in Cursor or Claude Desktop, add an mcpServers entry with command "npx", args ["-y", "@finedata/mcp-server"] and FINEDATA_API_KEY in env. For remote-capable clients, add https://mcp.finedata.ai/mcp using their connector UI or config. Authenticate with OAuth where supported, or an Authorization: Bearer API-key header. Remote configuration varies by client.

Do I need an API key, or does OAuth replace it?

Either works. An API key is simplest for a server-side agent you control. OAuth suits a client you hand to someone else: the user signs in to their own FineData account, approves the scopes on a consent screen, and usage bills to their account rather than yours. Nothing is copied or pasted, and the connection can be revoked later.

What can an agent do with an OAuth connection to my account?

The hosted connection requests scrape:write for scrapes and cancellation, jobs:read for polling, and usage:read for consumption. The token cannot create API keys, change plans, manage wallet operations or read profile data. Scrapes still incur normal usage charges. Access tokens last an hour, with refresh for up to 30 days; revoke access from the dashboard or by changing your password.

How much does an MCP scrape cost?

A default fetch uses 3 tokens before retries, extraction or rendering. Failed scrapes are not billed. Successful scrape results report actual tokens used; get_usage reports current-period consumption, not a balance or plan limit. See the API token-pricing section for mode costs and the billing dashboard for account limits.

Which modes should an agent try first?

Start with a default request. Use use_js_render when the content needs JavaScript. For a failed scrape, inspect the failure hint before choosing another engine or routing option. Fix authentication, billing and rate-limit errors directly rather than escalating rendering. No mode guarantees access to a particular site.

Is the MCP server open source, and where is it published?

The server is published as finedata-mcp on PyPI and @finedata/mcp-server on npm, and is listed in the MCP Registry. The remote endpoint publishes OAuth Authorization Server Metadata and Protected Resource Metadata, so a compliant client can discover and connect to it without any hand-written configuration.

Agents act quickly and at volume, which makes the rules matter more, not less. Collect only what you have a lawful basis to collect, respect each site's terms and robots directives, and keep request rates civil. Our Acceptable Use Policy sets the boundaries, and we make no representation that scraping any particular site is permissible.