Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
130 changes: 109 additions & 21 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,6 @@
# Context Dev Ruby API library

The Context Dev Ruby library provides convenient access to the Context Dev REST API from any Ruby 3.2.0+ application. It ships with comprehensive types & docstrings in Yard, RBS, and RBI – [see below](https://github.com/context-dot-dev/context-ruby-sdk#Sorbet) for usage with Sorbet. The standard library's `net/http` is used as the HTTP transport, with connection pooling via the `connection_pool` gem.

It is generated with [Stainless](https://www.stainless.com/).
Context.dev is a web scraping API for AI agents and LLMs. This SDK turns any URL into clean, LLM-ready markdown, crawls whole sites, searches the web, takes screenshots and extracts structured JSON against a schema you define, all with one API key. Proxies, JavaScript rendering and anti-bot handling run on Context.dev's side, so there is no headless browser to host.

## Documentation

Expand All @@ -24,26 +22,114 @@ gem "context.dev", "~> 2.24.0"

## Usage

Set `CONTEXT_DEV_API_KEY` to your API key; the client reads it automatically.

### Scrape markdown and HTML

```ruby
require "bundler/setup"
require "context_dev"

context_dev = ContextDev::Client.new(
api_key: ENV["CONTEXT_DEV_API_KEY"] # This is the default and can be omitted
context_dev = ContextDev::Client.new

page = context_dev.web.scrape(
url: "https://example.com",
formats: {markdown: true, html: true}
)

puts(page.markdown.data)
puts(page.html.data)
```

### Extract structured JSON

```ruby
require "bundler/setup"
require "context_dev"

context_dev = ContextDev::Client.new

page = context_dev.web.scrape(
url: "https://example.com",
formats: {json: true},
json_params: {
schema: {
type: "object",
properties: {
title: {type: ["string", "null"]},
description: {type: ["string", "null"]}
},
required: ["title", "description"],
additionalProperties: false
}
}
)

puts(page.json.data)
```

### Extract relevant highlights

Return the passages that answer a question about the page.

```ruby
require "bundler/setup"
require "context_dev"

context_dev = ContextDev::Client.new

page = context_dev.web.scrape(
url: "https://example.com",
formats: {highlights: true},
highlights_params: {query: "What is this domain used for?"}
)

brand = context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"})
puts(page.highlights.data)
```

### Take a screenshot

The screenshot is returned as a base64 image data URL.

```ruby
require "bundler/setup"
require "context_dev"

context_dev = ContextDev::Client.new

page = context_dev.web.scrape(
url: "https://example.com",
formats: {screenshot: true}
)

puts(brand.request_id)
puts(page.screenshot.data)
```

## What you can do

| Task | Method |
| --- | --- |
| Scrape a URL to markdown, HTML, JSON, highlights or a screenshot | `context_dev.web.scrape` |
| Crawl a site and get every page as markdown | `context_dev.web.web_crawl_md` |
| Map every URL on a domain | `context_dev.web.map_urls` |
| Search the web | `context_dev.web.search` |
| Take a screenshot of a page | `context_dev.web.screenshot` |
| Parse PDFs and documents | `context_dev.parse.handle` |
| Run thousands of URLs as a batch | `context_dev.batch.submit` |
| Watch a page for changes | `context_dev.monitors.create` |
| Look up a company's logo, colors and brand data | `context_dev.brand.retrieve` |

## Use it from an AI agent

Context.dev also ships as a plugin for [Claude](https://github.com/context-dot-dev/claude-plugin), [Cursor](https://github.com/context-dot-dev/cursor-plugin) and [Gemini CLI](https://github.com/context-dot-dev/gemini-cli-context), and as tools for [LangChain](https://github.com/context-dot-dev/langchain-context) and [Haystack](https://github.com/context-dot-dev/context-haystack).

### Handling errors

When the library is unable to connect to the API, or if the API returns a non-success status code (i.e., 4xx or 5xx response), a subclass of `ContextDev::Errors::APIError` will be thrown:

```ruby
begin
brand = context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"})
page = context_dev.web.scrape(url: "https://example.com", formats: {markdown: true})
rescue ContextDev::Errors::APIConnectionError => e
puts("The server could not be reached")
puts(e.cause) # an underlying Exception, likely raised within `net/http`
Expand Down Expand Up @@ -86,8 +172,8 @@ context_dev = ContextDev::Client.new(
)

# Or, configure per-request:
context_dev.brand.retrieve(
body: {domain: "stripe.com", type: "by_domain"},
context_dev.web.scrape(
url: "https://example.com", formats: {markdown: true},
request_options: {max_retries: 5}
)
```
Expand All @@ -103,7 +189,7 @@ context_dev = ContextDev::Client.new(
)

# Or, configure per-request:
context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"}, request_options: {timeout: 5})
context_dev.web.scrape(url: "https://example.com", formats: {markdown: true}, request_options: {timeout: 5})
```

On timeout, `ContextDev::Errors::APITimeoutError` is raised.
Expand Down Expand Up @@ -133,17 +219,17 @@ You can send undocumented parameters to any endpoint, and read undocumented resp
Note: the `extra_` parameters of the same name overrides the documented parameters.

```ruby
brand =
context_dev.brand.retrieve(
body: {domain: "stripe.com", type: "by_domain"},
page =
context_dev.web.scrape(
url: "https://example.com", formats: {markdown: true},
request_options: {
extra_query: {my_query_parameter: value},
extra_body: {my_body_parameter: value},
extra_headers: {"my-header": value}
}
)

puts(brand[:my_undocumented_property])
puts(page[:my_undocumented_property])
```

#### Undocumented request params
Expand Down Expand Up @@ -181,22 +267,24 @@ This library provides comprehensive [RBI](https://sorbet.org/docs/rbi) definitio
You can provide typesafe request parameters like so:

```ruby
context_dev.brand.retrieve(
body: ContextDev::BrandRetrieveParams::Body::ByDomain.new(domain: "stripe.com")
context_dev.web.scrape(
url: "https://example.com",
formats: ContextDev::WebScrapeParams::Formats.new(markdown: true)
)
```

Or, equivalently:

```ruby
# Hashes work, but are not typesafe:
context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"})
context_dev.web.scrape(url: "https://example.com", formats: {markdown: true})

# You can also splat a full Params class:
params = ContextDev::BrandRetrieveParams.new(
body: ContextDev::BrandRetrieveParams::Body::ByDomain.new(domain: "stripe.com")
params = ContextDev::WebScrapeParams.new(
url: "https://example.com",
formats: ContextDev::WebScrapeParams::Formats.new(markdown: true)
)
context_dev.brand.retrieve(**params)
context_dev.web.scrape(**params)
```

### Enums
Expand Down
Loading