Packages

Profile Runner

Extract website data using reusable YAML or JSON profiles.

@scrape-admin/profile-runner executes scraping profiles that describe how to load a page, extract values, and transform the results. Use it from an application with the package API, or run profiles directly with its CLI.

Profile structure

A profile can be provided as a JavaScript object or as YAML/JSON text. It can define a URL, request headers, settings, actions, and elements to extract:

url: https://example.com
headers:
  User-Agent: Custom UA
settings:
  timeout: 5000
elements:
  - type: attribute
    name: title
    selector: h1.title
    resultSelector: 0.text
  - type: variable
    name: categories
    selector: ul.categories > li
    resultSelector: '*.text'

Element selectors can select DOM nodes or contain custom code. Attribute elements extract page data, while variable elements store intermediate values for other extractors. resultSelector selects properties from matched elements, such as 0.text for the first match's text or *.text for all matched texts.

Actions and transformations

Actions run before or after extraction. Built-in browser actions include waitForElement, clickElement, fillField, pressKey, waitForTimeout, and loop. Code actions can modify the extraction context.

Transformers process an extracted value in order. Built-ins include convert.json, number.toFloat, number.toInt, number, text.replace, text.toString, and custom.

elements:
  - type: attribute
    name: price
    selector: .price
    resultSelector: 0.text
    transform:
      - callback: text.replace
        params: ['/[^0-9.]/g', '']
      - callback: number.toFloat

Browser mode

Use runProfile for browser-based scraping. It can create and manage a browser page or use an existing Puppeteer browser or page:

import { runProfile } from '@scrape-admin/profile-runner'

const result = await runProfile({
  profile: {
    elements: [
      { type: 'attribute', name: 'title', selector: 'h1', resultSelector: '0.text' },
    ],
  },
  urlOverride: 'https://example.com',
})

Pass browser or page to reuse an existing Puppeteer instance. Profiles can also configure request headers, block requests, set a proxy server, and define a timeout.

Browser-free mode

For static or server-rendered pages, runProfileWithCheerio uses fetch and Cheerio rather than launching a browser:

import { runProfileWithCheerio } from '@scrape-admin/profile-runner/cheerio'

const result = await runProfileWithCheerio({
  profile: {
    url: 'https://example.com',
    elements: [
      { type: 'attribute', name: 'title', selector: 'h1', resultSelector: '0.text' },
    ],
  },
  fetchOptions: { headers: { 'User-Agent': 'Custom UA' } },
})

This mode supports extraction, variables, transformations, custom code, and waitForTimeout and code actions. It does not support browser interactions such as clicking or filling fields. Request-blocking and the profile's proxyServer setting are ignored; configure a proxy-aware fetch implementation through fetchOptions instead.

CLI

The profile-runner command reads a profile file or accepts YAML/JSON from standard input:

pnpm --filter @scrape-admin/profile-runner run:profile --profile ./profile.yaml --url https://example.com
cat profile.yaml | pnpm --filter @scrape-admin/profile-runner run:profile --url https://example.com

Use --ws to connect to a remote browser WebSocket endpoint, --proxy to specify a proxy, --format to select JSON or YAML output, and --pretty to format JSON output. --url overrides the URL in the profile.

For the full profile reference and additional examples, see the profile-runner README.

Copyright © 2026