Website scraping API for rendered HTML, clean text and published emails
Send a public URL and get back the page's final HTML, the text a visitor sees, or the email addresses the site publishes. JavaScript is rendered when the page needs it. Every response is JSON with a fixed credit price.
GET/v1/websites/text?url=https://example.com/contactGet API keyRead the docs100 free creditsNo credit card required
Endpoints
Three endpoints, one url parameter
Each endpoint takes the same url query parameter and returns the final URL it read. Pick the output you need: raw HTML, visible text, or email addresses.
Get website HTML
Returns the full HTML of the page after redirects, rendered when the page needs JavaScript. Use it when you want to run your own parser, keep a copy, or compare versions over time.
- Path
/v1/websites/content- Returns
url · title · html- Cost
2 credits
Get website text
Returns the text a visitor can read, with markup, scripts and styles removed. Use it for search indexes, classification, summaries and LLM prompts.
- Path
/v1/websites/text- Returns
url · title · text- Cost
2 credits
Find published emails
Finds email addresses the site publishes. It reads up to eight pages and stops at the first page that has any. Each address comes with the page it was found on, so you can show where it came from.
- Path
/v1/websites/emails- Returns
emails[] · pagesChecked · stopReason- Cost
5 credits
Email search
How the email search works
One request to /v1/websites/emails does a bounded search of one site. It never runs longer than 8 pages, and it tells you why it stopped.
Start at your URL
The search opens the page you send. A contact or about page is the fastest start; a homepage works too.
Read up to 8 pages
It follows the site's own links to more pages of the same site when the first page has no addresses.
Stop at the first page with addresses
As soon as a page publishes email addresses, the search ends and returns them with the page each one came from.
Return up to 100 addresses
If the page had more, truncated is true. pagesChecked and stopReason tell you how far the search went.
| stopReason | What it means | What to do |
|---|---|---|
found | A page with addresses was reached. | Use the results. |
exhausted | The site ran out of linked pages before 8 were read, and none had addresses. | There is nothing published here. Try the Email Finder API with a name and the domain. |
page_limit | 8 pages were read and none had addresses. | Send a direct contact or about page URL instead of the homepage. |
time_limit | The search hit its time budget before it finished. | Retry with a more specific URL. |
page_errors | Pages could not be loaded, so the search could not continue. | Check that the site is reachable, then retry. |
Response
What each response contains
Every endpoint returns a { data, meta } object, and meta.creditCost tells you what the request cost. An email search also reports how many pages it read and why it stopped, so you can handle empty results in code instead of guessing.
url- The final URL the page was read from, after redirects.
title- The page title, or null when the page has none. Up to 4,096 characters.
html / text- The full page HTML, or its visible text. Up to 2 MB.
emails[]- Each address and the page it was published on, as email and sourceUrl.
pagesChecked- How many pages the search read, from 1 to 8.
truncated- True when the site published more than the 100 addresses returned.
stopReason- Why the search ended:
foundexhaustedpage_limittime_limitpage_errors meta.creditCost- Credits charged for this request.
curl -G "https://api.envoapi.com/v1/websites/content" \ -H "Authorization: Bearer $ENVO_API_KEY" \ --data-urlencode "url=https://example.com/contact"{ "data": { "url": "https://example.com/contact", "title": "Contact — Example", "html": "<!doctype html><html lang=\"en\"><head><title>Contact — Example…" }, "meta": { "creditCost": 2 }}curl -G "https://api.envoapi.com/v1/websites/text" \ -H "Authorization: Bearer $ENVO_API_KEY" \ --data-urlencode "url=https://example.com/contact"{ "data": { "url": "https://example.com/contact", "title": "Contact — Example", "text": "Contact\n\nWrite to us and we reply within one business day…" }, "meta": { "creditCost": 2 }}curl -G "https://api.envoapi.com/v1/websites/emails" \ -H "Authorization: Bearer $ENVO_API_KEY" \ --data-urlencode "url=https://example.com/contact"{ "data": { "url": "https://example.com/contact", "emails": [ { "email": "[email protected]", "sourceUrl": "https://example.com/contact" }, { "email": "[email protected]", "sourceUrl": "https://example.com/contact" } ], "pagesChecked": 1, "truncated": false, "stopReason": "found" }, "meta": { "creditCost": 5 }}Example responses are shortened. A real html or text field holds the whole page, up to 2 MB.
HTML or text
Which output to pick
Both endpoints run the same fetch, cost the same 2 credits, and return the same url and title. The difference is what you do with the page next.
/v1/websites/contentPick HTML when you parse, diff or archive
You get the page as a browser saw it, with every tag, link, meta tag and script in place. Run your own parser, keep a copy, or compare two fetches.
- Links, meta tags and structured data
- Version diffs of pricing or terms pages
- Archives you can re-parse later
/v1/websites/textPick text when you index, classify or prompt
You get what a visitor can read, with markup, scripts and styles removed. It is smaller, cleaner, and ready for a search index or a language model.
- Embeddings and retrieval indexes
- Industry and product classification
- Summaries and LLM context
Use cases
What teams build with it
/v1/websites/emailsFill in contact emails for CRM accounts
Take the domain you already have for an account and pull the email addresses the company publishes on its own site. Keep the source page next to each address so your team can see where it came from.
/v1/websites/textFeed web pages to LLMs and RAG pipelines
Get the readable text of a page without navigation, scripts and markup. Chunk it, embed it, or pass it straight into a prompt. JavaScript pages are rendered for you.
/v1/websites/textClassify companies from their own websites
Read the homepage, product and about pages as plain text, then tag industry, products and target market with your own model or rules.
/v1/websites/contentTrack changes on pricing, terms and changelog pages
Fetch the final HTML on a schedule, store it, and diff versions with your own parser. You get the page as a browser saw it, not the raw source.
/v1/websites/textQualify inbound signups
When a signup arrives with a company domain, read the homepage text to learn what the company does before you route the lead.
/v1/websites/contentReview a vendor or customer domain
Before you onboard an account, fetch its homepage to confirm the domain has a working site and see what it shows. Use the HTML when you need links, meta tags or structured data.
Built for production
What you can rely on
JavaScript handled for you
Static pages are read as served. Pages that build their content with JavaScript are rendered first. You never pick a mode.
Public web only
HTTP and HTTPS on ports 80 and 443. Private networks, localhost and internal hostnames are refused, so the API cannot be pointed at your own infrastructure.
Treat the output as untrusted
The HTML and text come from someone else's website. Sanitize it before you render it in a browser or hand it to a tool that runs code.
Fixed price per request
2 credits for HTML or text, 5 for an email search. The price is the same whether a search finds 100 addresses or none, so costs are easy to forecast.
Which URLs are accepted
The same check runs before any request is charged. A URL that fails it returns 400 and costs nothing.
Accepted
- HTTP or HTTPS, on the default port or an explicit 80 or 443
- Up to 2,048 characters
- Any public hostname with a dot in it, or a public IP address
- Query strings. A fragment after # is ignored
Refused
- A username or password in the URL
- Any port other than 80 or 443
- localhost, and hostnames ending in .local, .internal, .lan, .home, .test or .invalid
- Private, loopback, link-local and other reserved IP ranges, including hosts that resolve to one
FAQ
Frequently asked questions
What does the Website Scraping API return?
Three endpoints share one url parameter. /v1/websites/content returns the final page HTML, /v1/websites/text returns the visible text, and both include the final URL and the page title. /v1/websites/emails returns the email addresses a site publishes, each with the page it was found on.
Is this a web scraping API or a crawler?
The HTML and text endpoints read one page per request. The email search reads up to 8 pages of one site and stops at the first page with addresses. There is no site-wide crawl, so you always know the most a request can do.
Does it render JavaScript?
When the page needs it. Static pages are read as served. Pages that build their content with JavaScript are rendered before the HTML or text is extracted, so you never choose a rendering mode yourself.
Which URLs can I send?
Public HTTP or HTTPS URLs on ports 80 and 443, up to 2,048 characters, without a username or password. Private, internal and localhost addresses are refused.
Can I fetch pages behind a login?
No. The API reads public pages only. URLs that carry a username or password are rejected, and no cookies or headers of yours are forwarded to the site.
Does it follow redirects?
Yes. The url field in every response is the final address the page was read from, so you can store it next to the address you sent.
How does the email search work?
It starts at the URL you send, reads up to 8 pages of the site, and stops at the first page that publishes email addresses. It returns up to 100 addresses with their source pages, how many pages it read, and a stopReason that says why it ended: found, exhausted, page_limit, time_limit or page_errors.
Are the emails verified?
No. The search reports addresses a site publishes. It does not check who owns them or whether they accept mail. Run an address through the Email Verification API before you send to it.
Which email should I use from the results?
Each address comes with a sourceUrl. Prefer addresses found on contact or about pages, then verify them before sending. If the site publishes nothing, the Email Finder API can look up a person's work email from their name and domain.
How large can a response be?
HTML and text are returned up to 2 MB. The title is up to 4,096 characters. An email search returns up to 100 addresses and sets truncated to true when the site had more.
What happens when the site is slow or down?
The request fails with a clear status instead of a partial page. You get 504 when the site times out, 502 when it returns an invalid response, and 503 with error.retryable when you should retry. Every response carries X-Request-Id for support and X-RateLimit-Remaining for pacing.
How many credits does a request cost?
HTML and text cost 2 credits each, and an email search costs 5. Email searches are charged whether they find addresses, find none, or stop early. New accounts start with 100 free credits, enough for 50 page fetches or 20 email searches.
Is the returned HTML safe to display?
Treat it as untrusted. The HTML and text come from a third-party site, so sanitize them before rendering them in a browser or passing them to a tool that executes content.
Try it on a page you already know
Sign up, copy your API key, and send your first URL. 100 free credits cover 50 page fetches or 20 email searches.