From e74b1f9c98402dd743b70928fbcc4629fa0f4966 Mon Sep 17 00:00:00 2001 From: Yahia Bakour Date: Thu, 8 Oct 2026 12:48:22 -0400 Subject: [PATCH 1/3] docs: lead README with Context.dev capabilities --- README.md | 47 ++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 46 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index d993fe77..42ba0496 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # Context Dev Ruby API library -The Context Dev Ruby library provides convenient access to the Context Dev REST API from any Ruby 3.2.0+ application. It ships with comprehensive types & docstrings in Yard, RBS, and RBI – [see below](https://github.com/context-dot-dev/context-ruby-sdk#Sorbet) for usage with Sorbet. The standard library's `net/http` is used as the HTTP transport, with connection pooling via the `connection_pool` gem. +Context.dev is a web scraping API for AI agents and LLMs. This SDK turns any URL into clean, LLM-ready markdown, crawls whole sites, searches the web, takes screenshots and extracts structured JSON against a schema you define, all with one API key. Proxies, JavaScript rendering and anti-bot handling run on Context.dev's side, so there is no headless browser to host. It is generated with [Stainless](https://www.stainless.com/). @@ -37,6 +37,51 @@ brand = context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain puts(brand.request_id) ``` +### Extract structured JSON + +```ruby +require "bundler/setup" +require "context_dev" + +context_dev = ContextDev::Client.new + +page = context_dev.web.scrape( + url: "https://example.com", + formats: {json: true}, + json_params: { + schema: { + type: "object", + properties: { + title: {type: ["string", "null"]}, + description: {type: ["string", "null"]} + }, + required: ["title", "description"], + additionalProperties: false + } + } +) + +puts(page.json.data) +``` + +## What you can do + +| Task | Method | +| --- | --- | +| Scrape a URL to markdown, HTML, JSON or a screenshot | `context_dev.web.scrape` | +| Crawl a site and get every page as markdown | `context_dev.web.web_crawl_md` | +| Map every URL on a domain | `context_dev.web.map_urls` | +| Search the web | `context_dev.web.search` | +| Take a screenshot of a page | `context_dev.web.screenshot` | +| Parse PDFs and documents | `context_dev.parse.handle` | +| Run thousands of URLs as a batch | `context_dev.batch.submit` | +| Watch a page for changes | `context_dev.monitors.create` | +| Look up a company's logo, colors and brand data | `context_dev.brand.retrieve` | + +## Use it from an AI agent + +Context.dev also ships as a plugin for [Claude](https://github.com/context-dot-dev/claude-plugin), [Cursor](https://github.com/context-dot-dev/cursor-plugin) and [Gemini CLI](https://github.com/context-dot-dev/gemini-cli-context), and as tools for [LangChain](https://github.com/context-dot-dev/langchain-context) and [Haystack](https://github.com/context-dot-dev/context-haystack). + ### Handling errors When the library is unable to connect to the API, or if the API returns a non-success status code (i.e., 4xx or 5xx response), a subclass of `ContextDev::Errors::APIError` will be thrown: From c9d5c4ff4523343ffa7a2c29ab7789f6d3cb81fe Mon Sep 17 00:00:00 2001 From: Yahia Bakour Date: Thu, 8 Oct 2026 12:51:20 -0400 Subject: [PATCH 2/3] docs: remove Stainless README attribution --- README.md | 2 -- 1 file changed, 2 deletions(-) diff --git a/README.md b/README.md index 42ba0496..87e6d85a 100644 --- a/README.md +++ b/README.md @@ -2,8 +2,6 @@ Context.dev is a web scraping API for AI agents and LLMs. This SDK turns any URL into clean, LLM-ready markdown, crawls whole sites, searches the web, takes screenshots and extracts structured JSON against a schema you define, all with one API key. Proxies, JavaScript rendering and anti-bot handling run on Context.dev's side, so there is no headless browser to host. -It is generated with [Stainless](https://www.stainless.com/). - ## Documentation Documentation for releases of this gem can be found [on RubyDoc](https://gemdocs.org/gems/context.dev). From e04afa48893c24235a6c70d0e07af1eeda9224c4 Mon Sep 17 00:00:00 2001 From: Yahia Bakour Date: Thu, 8 Oct 2026 13:03:56 -0400 Subject: [PATCH 3/3] docs: demonstrate scraping formats throughout README --- README.md | 85 ++++++++++++++++++++++++++++++++++++++++++------------- 1 file changed, 65 insertions(+), 20 deletions(-) diff --git a/README.md b/README.md index 87e6d85a..16dbc0e8 100644 --- a/README.md +++ b/README.md @@ -22,17 +22,23 @@ gem "context.dev", "~> 2.24.0" ## Usage +Set `CONTEXT_DEV_API_KEY` to your API key; the client reads it automatically. + +### Scrape markdown and HTML + ```ruby require "bundler/setup" require "context_dev" -context_dev = ContextDev::Client.new( - api_key: ENV["CONTEXT_DEV_API_KEY"] # This is the default and can be omitted -) +context_dev = ContextDev::Client.new -brand = context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"}) +page = context_dev.web.scrape( + url: "https://example.com", + formats: {markdown: true, html: true} +) -puts(brand.request_id) +puts(page.markdown.data) +puts(page.html.data) ``` ### Extract structured JSON @@ -62,11 +68,48 @@ page = context_dev.web.scrape( puts(page.json.data) ``` +### Extract relevant highlights + +Return the passages that answer a question about the page. + +```ruby +require "bundler/setup" +require "context_dev" + +context_dev = ContextDev::Client.new + +page = context_dev.web.scrape( + url: "https://example.com", + formats: {highlights: true}, + highlights_params: {query: "What is this domain used for?"} +) + +puts(page.highlights.data) +``` + +### Take a screenshot + +The screenshot is returned as a base64 image data URL. + +```ruby +require "bundler/setup" +require "context_dev" + +context_dev = ContextDev::Client.new + +page = context_dev.web.scrape( + url: "https://example.com", + formats: {screenshot: true} +) + +puts(page.screenshot.data) +``` + ## What you can do | Task | Method | | --- | --- | -| Scrape a URL to markdown, HTML, JSON or a screenshot | `context_dev.web.scrape` | +| Scrape a URL to markdown, HTML, JSON, highlights or a screenshot | `context_dev.web.scrape` | | Crawl a site and get every page as markdown | `context_dev.web.web_crawl_md` | | Map every URL on a domain | `context_dev.web.map_urls` | | Search the web | `context_dev.web.search` | @@ -86,7 +129,7 @@ When the library is unable to connect to the API, or if the API returns a non-su ```ruby begin - brand = context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"}) + page = context_dev.web.scrape(url: "https://example.com", formats: {markdown: true}) rescue ContextDev::Errors::APIConnectionError => e puts("The server could not be reached") puts(e.cause) # an underlying Exception, likely raised within `net/http` @@ -129,8 +172,8 @@ context_dev = ContextDev::Client.new( ) # Or, configure per-request: -context_dev.brand.retrieve( - body: {domain: "stripe.com", type: "by_domain"}, +context_dev.web.scrape( + url: "https://example.com", formats: {markdown: true}, request_options: {max_retries: 5} ) ``` @@ -146,7 +189,7 @@ context_dev = ContextDev::Client.new( ) # Or, configure per-request: -context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"}, request_options: {timeout: 5}) +context_dev.web.scrape(url: "https://example.com", formats: {markdown: true}, request_options: {timeout: 5}) ``` On timeout, `ContextDev::Errors::APITimeoutError` is raised. @@ -176,9 +219,9 @@ You can send undocumented parameters to any endpoint, and read undocumented resp Note: the `extra_` parameters of the same name overrides the documented parameters. ```ruby -brand = - context_dev.brand.retrieve( - body: {domain: "stripe.com", type: "by_domain"}, +page = + context_dev.web.scrape( + url: "https://example.com", formats: {markdown: true}, request_options: { extra_query: {my_query_parameter: value}, extra_body: {my_body_parameter: value}, @@ -186,7 +229,7 @@ brand = } ) -puts(brand[:my_undocumented_property]) +puts(page[:my_undocumented_property]) ``` #### Undocumented request params @@ -224,8 +267,9 @@ This library provides comprehensive [RBI](https://sorbet.org/docs/rbi) definitio You can provide typesafe request parameters like so: ```ruby -context_dev.brand.retrieve( - body: ContextDev::BrandRetrieveParams::Body::ByDomain.new(domain: "stripe.com") +context_dev.web.scrape( + url: "https://example.com", + formats: ContextDev::WebScrapeParams::Formats.new(markdown: true) ) ``` @@ -233,13 +277,14 @@ Or, equivalently: ```ruby # Hashes work, but are not typesafe: -context_dev.brand.retrieve(body: {domain: "stripe.com", type: "by_domain"}) +context_dev.web.scrape(url: "https://example.com", formats: {markdown: true}) # You can also splat a full Params class: -params = ContextDev::BrandRetrieveParams.new( - body: ContextDev::BrandRetrieveParams::Body::ByDomain.new(domain: "stripe.com") +params = ContextDev::WebScrapeParams.new( + url: "https://example.com", + formats: ContextDev::WebScrapeParams::Formats.new(markdown: true) ) -context_dev.brand.retrieve(**params) +context_dev.web.scrape(**params) ``` ### Enums