Profile Runner
@scrape-admin/profile-runner executes scraping profiles that describe how to load a page, extract values, and transform
the results. Use it from an application with the package API, or run profiles directly with its CLI.
Profile structure
A profile can be provided as a JavaScript object or as YAML/JSON text. It can define a URL, request headers, settings, actions, and elements to extract:
url: https://example.com
headers:
User-Agent: Custom UA
settings:
timeout: 5000
elements:
- type: attribute
name: title
selector: h1.title
resultSelector: 0.text
- type: variable
name: categories
selector: ul.categories > li
resultSelector: '*.text'
Element selectors can select DOM nodes or contain custom code. Attribute elements extract page data, while variable
elements store intermediate values for other extractors. resultSelector selects properties from matched elements,
such as 0.text for the first match's text or *.text for all matched texts.
Actions and transformations
Actions run before or after extraction. Built-in browser actions include waitForElement, clickElement, fillField,
pressKey, waitForTimeout, and loop. Code actions can modify the extraction context.
Transformers process an extracted value in order. Built-ins include convert.json, number.toFloat, number.toInt,
number, text.replace, text.toString, and custom.
elements:
- type: attribute
name: price
selector: .price
resultSelector: 0.text
transform:
- callback: text.replace
params: ['/[^0-9.]/g', '']
- callback: number.toFloat
Browser mode
Use runProfile for browser-based scraping. It can create and manage a browser page or use an existing Puppeteer browser
or page:
import { runProfile } from '@scrape-admin/profile-runner'
const result = await runProfile({
profile: {
elements: [
{ type: 'attribute', name: 'title', selector: 'h1', resultSelector: '0.text' },
],
},
urlOverride: 'https://example.com',
})
Pass browser or page to reuse an existing Puppeteer instance. Profiles can also configure request headers, block
requests, set a proxy server, and define a timeout.
Browser-free mode
For static or server-rendered pages, runProfileWithCheerio uses fetch and Cheerio rather than launching a browser:
import { runProfileWithCheerio } from '@scrape-admin/profile-runner/cheerio'
const result = await runProfileWithCheerio({
profile: {
url: 'https://example.com',
elements: [
{ type: 'attribute', name: 'title', selector: 'h1', resultSelector: '0.text' },
],
},
fetchOptions: { headers: { 'User-Agent': 'Custom UA' } },
})
This mode supports extraction, variables, transformations, custom code, and waitForTimeout and code actions. It
does not support browser interactions such as clicking or filling fields. Request-blocking and the profile's
proxyServer setting are ignored; configure a proxy-aware fetch implementation through fetchOptions instead.
CLI
The profile-runner command reads a profile file or accepts YAML/JSON from standard input:
pnpm --filter @scrape-admin/profile-runner run:profile --profile ./profile.yaml --url https://example.com
cat profile.yaml | pnpm --filter @scrape-admin/profile-runner run:profile --url https://example.com
Use --ws to connect to a remote browser WebSocket endpoint, --proxy to specify a proxy, --format to select JSON or
YAML output, and --pretty to format JSON output. --url overrides the URL in the profile.
For the full profile reference and additional examples, see the profile-runner README.