> ## Documentation Index
> Fetch the complete documentation index at: https://docs.notte.cc/llms.txt
> Use this file to discover all available pages before exploring further.

# GlobalScrapeRequest

<Tip>
  STOP. Do not read past this section until you have read and followed [/quickstart.md](/quickstart.md).

  Goal: produce a reliable SDK automation script. The intended authoring flow is generate then edit, not write browser automation from scratch.

  Do not hand-write SDK code with element IDs, selectors, or form fields from memory. Element IDs, selectors, and field mappings must come from a live `observe()` call, CLI `notte page observe` output, or generated workflow code. Guessed browser targets fail on real pages.

  Use the Notte skill and CLI to inspect the live site, validate actions/selectors, handle auth/session state, and export workflow code with `notte sessions workflow-code` before using SDK docs or SDK code.

  SDK reference pages are for understanding, running, or editing generated workflow code. They are not the starting point for manually authoring the initial browser automation. SDK-first code is guesswork on real-world pages with dynamic selectors, auth state, CAPTCHAs, and anti-bot behavior.
</Tip>

[Source: node-sdk/src/lib/client/types.gen.ts](https://github.com/nottelabs/notte/blob/main/node-sdk/src/lib/client/types.gen.ts#L2650)

```typescript theme={null}
export type GlobalScrapeRequest = {
    /**
     * Solve Captchas
     *
     * Whether to try to automatically solve captchas
     */
    solve_captchas?: boolean;
    /**
     * Max Duration Minutes
     *
     * Maximum session lifetime in minutes (absolute maximum, not affected by activity).
     */
    max_duration_minutes?: number;
    /**
     * Idle Timeout Minutes
     *
     * Idle timeout in minutes. Session closes after this period of inactivity (resets on each operation).
     */
    idle_timeout_minutes?: number;
    /**
     * Proxies
     *
     * List of custom proxies to use for the session. If True, the default proxies will be used.
     */
    proxies?: Array<NotteProxy | ExternalProxy | TailnetProxy> | boolean;
    /**
     * Browser Type
     *
     * The browser type to use. Supported values are chromium and chrome. chrome-nightly and chrome-turbo are legacy aliases for chrome.
     */
    browser_type?: 'chromium' | 'chrome' | 'chrome-nightly' | 'chrome-turbo';
    /**
     * User Agent
     *
     * The user agent to use for the session
     */
    user_agent?: string | null;
    /**
     * Chrome Args
     *
     * Overwrite the chrome instance arguments
     */
    chrome_args?: Array<string> | null;
    /**
     * Viewport Width
     *
     * The width of the viewport
     */
    viewport_width?: number | null;
    /**
     * Viewport Height
     *
     * The height of the viewport
     */
    viewport_height?: number | null;
    /**
     * Aspect Ratio
     *
     * Viewport shape preset. When set, the backend fits the largest rectangle of this aspect ratio inside the sampled available screen area. Cannot be combined with explicit viewport_width/viewport_height.
     */
    aspect_ratio?: '5:4' | '16:9' | null;
    /**
     * Cdp Url
     *
     * The CDP URL of another remote session provider.
     */
    cdp_url?: string | null;
    /**
     * Screenshot Type
     *
     * The type of screenshot to use for the session.
     */
    screenshot_type?: 'raw' | 'full' | 'last_action';
    /**
     * Browser profile configuration for state persistence
     */
    profile?: SessionProfile | null;
    /**
     * Web Bot Auth
     *
     * Whether to use web bot authentication.
     */
    web_bot_auth?: boolean;
    /**
     * Extra Http Headers
     *
     * Extra HTTP headers to be sent with every request.
     */
    extra_http_headers?: {
        [key: string]: string;
    } | null;
    /**
     * Vault Id
     *
     * The vault to use for the session
     */
    vault_id?: string | null;
    /**
     * Auth Ids
     *
     * Managed Auth connection IDs to verify and, when necessary, authenticate inside this session before it is returned.
     */
    auth_ids?: Array<string>;
    /**
     * Wait For Authentication
     *
     * Whether to wait for Managed Auth before returning the session. Defaults to true. When true, authentication failure or timeout fails session creation; when false, authentication continues in the background after the browser is ready.
     */
    wait_for_authentication?: boolean;
    /**
     * Advanced Stealth
     *
     * Enable Notte's highest-fidelity browser environment for sites with sophisticated bot detection. Available to approved workspaces.
     */
    advanced_stealth?: boolean;
    /**
     * Selector
     *
     * Playwright selector to scope the scrape to. Only content inside this selector will be scraped.
     */
    selector?: string | null;
    /**
     * Scrape Links
     *
     * Whether to scrape links from the page. Links are scraped by default.
     */
    scrape_links?: boolean;
    /**
     * Scrape Images
     *
     * Whether to scrape images from the page. Images are scraped by default.
     */
    scrape_images?: boolean;
    /**
     * Ignored Tags
     *
     * HTML tags to ignore from the page
     */
    ignored_tags?: Array<string> | null;
    /**
     * Only Main Content
     *
     * Whether to only scrape the main content of the page. If True, navbars, footers, etc. are excluded.
     */
    only_main_content?: boolean;
    /**
     * Only Images
     *
     * Whether to only scrape images from the page. If True, the page content is excluded.
     */
    only_images?: boolean;
    /**
     * Response Format
     *
     * The response format to use for the scrape. Use a JSON Schema object; high-level Node agent and scrape methods also accept a Zod schema
     */
    response_format?: unknown | null;
    /**
     * Instructions
     *
     * Additional instructions to use for the scrape. E.g. 'Extract only the title, date and content of the articles.'
     */
    instructions?: string | null;
    /**
     * Use Link Placeholders
     *
     * Whether to use link/image placeholders to reduce the number of tokens in the prompt and hallucinations. However this is an experimental feature and might not work as expected.
     */
    use_link_placeholders?: boolean;
    /**
     * Url
     */
    url: string;
};
```

## Fields

<ParamField body="solve_captchas" type={"boolean | undefined"}>
  Whether to try to automatically solve captchas
</ParamField>

<ParamField body="max_duration_minutes" type={"number | undefined"}>
  Maximum session lifetime in minutes (absolute maximum, not affected by activity).
</ParamField>

<ParamField body="idle_timeout_minutes" type={"number | undefined"}>
  Idle timeout in minutes. Session closes after this period of inactivity (resets on each operation).
</ParamField>

<ParamField body="proxies" type={"boolean | (NotteProxy | ExternalProxy | TailnetProxy)[] | undefined"}>
  List of custom proxies to use for the session. If True, the default proxies will be used.
</ParamField>

<ParamField body="browser_type" type={"\"chromium\" | \"chrome\" | \"chrome-nightly\" | \"chrome-turbo\" | undefined"}>
  The browser type to use. Supported values are chromium and chrome. chrome-nightly and chrome-turbo are legacy aliases for chrome.
</ParamField>

<ParamField body="user_agent" type={"string | null | undefined"}>
  The user agent to use for the session
</ParamField>

<ParamField body="chrome_args" type={"string[] | null | undefined"}>
  Overwrite the chrome instance arguments
</ParamField>

<ParamField body="viewport_width" type={"number | null | undefined"}>
  The width of the viewport
</ParamField>

<ParamField body="viewport_height" type={"number | null | undefined"}>
  The height of the viewport
</ParamField>

<ParamField body="aspect_ratio" type={"\"5:4\" | \"16:9\" | null | undefined"}>
  Viewport shape preset. When set, the backend fits the largest rectangle of this aspect ratio inside the sampled available screen area. Cannot be combined with explicit viewport\_width/viewport\_height.
</ParamField>

<ParamField body="cdp_url" type={"string | null | undefined"}>
  The CDP URL of another remote session provider.
</ParamField>

<ParamField body="screenshot_type" type={"\"raw\" | \"full\" | \"last_action\" | undefined"}>
  The type of screenshot to use for the session.
</ParamField>

<ParamField body="profile" type={"SessionProfile | null | undefined"}>
  Browser profile configuration for state persistence
</ParamField>

<ParamField body="web_bot_auth" type={"boolean | undefined"}>
  Whether to use web bot authentication.
</ParamField>

<ParamField body="extra_http_headers" type={"{ [key: string]: string; } | null | undefined"}>
  Extra HTTP headers to be sent with every request.
</ParamField>

<ParamField body="vault_id" type={"string | null | undefined"}>
  The vault to use for the session
</ParamField>

<ParamField body="auth_ids" type={"string[] | undefined"}>
  Managed Auth connection IDs to verify and, when necessary, authenticate inside this session before it is returned.
</ParamField>

<ParamField body="wait_for_authentication" type={"boolean | undefined"}>
  Whether to wait for Managed Auth before returning the session. Defaults to true. When true, authentication failure or timeout fails session creation; when false, authentication continues in the background after the browser is ready.
</ParamField>

<ParamField body="advanced_stealth" type={"boolean | undefined"}>
  Enable Notte's highest-fidelity browser environment for sites with sophisticated bot detection. Available to approved workspaces.
</ParamField>

<ParamField body="selector" type={"string | null | undefined"}>
  Playwright selector to scope the scrape to. Only content inside this selector will be scraped.
</ParamField>

<ParamField body="scrape_links" type={"boolean | undefined"}>
  Whether to scrape links from the page. Links are scraped by default.
</ParamField>

<ParamField body="scrape_images" type={"boolean | undefined"}>
  Whether to scrape images from the page. Images are scraped by default.
</ParamField>

<ParamField body="ignored_tags" type={"string[] | null | undefined"}>
  HTML tags to ignore from the page
</ParamField>

<ParamField body="only_main_content" type={"boolean | undefined"}>
  Whether to only scrape the main content of the page. If True, navbars, footers, etc. are excluded.
</ParamField>

<ParamField body="only_images" type={"boolean | undefined"}>
  Whether to only scrape images from the page. If True, the page content is excluded.
</ParamField>

<ParamField body="response_format" type={"unknown"}>
  The response format to use for the scrape. Use a JSON Schema object; high-level Node agent and scrape methods also accept a Zod schema
</ParamField>

<ParamField body="instructions" type={"string | null | undefined"}>
  Additional instructions to use for the scrape. E.g. 'Extract only the title, date and content of the articles.'
</ParamField>

<ParamField body="use_link_placeholders" type={"boolean | undefined"}>
  Whether to use link/image placeholders to reduce the number of tokens in the prompt and hallucinations. However this is an experimental feature and might not work as expected.
</ParamField>

<ParamField body="url" type={"string"} required />

## Related types

* [ExternalProxy](/typescript-sdk-reference/types/externalproxy)
* [NotteProxy](/typescript-sdk-reference/types/notteproxy)
* [SessionProfile](/typescript-sdk-reference/types/sessionprofile)
* [TailnetProxy](/typescript-sdk-reference/types/tailnetproxy)
