# Http-request node does not get all the page data (AKA How to scrape dynamic data from a web page)

**URL:** <https://discourse.nodered.org/t/http-request-node-does-not-get-all-the-page-data-aka-how-to-scrape-dynamic-data-from-a-web-page/69164>\
**Category:** FAQs\
**Created:** [17 October 2022 18:24 UTC](https://discourse.nodered.org/t/http-request-node-does-not-get-all-the-page-data-aka-how-to-scrape-dynamic-data-from-a-web-page/69164 "2022-10-17T18:24:44Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [17 October 2022 18:24 UTC](https://discourse.nodered.org/t/http-request-node-does-not-get-all-the-page-data-aka-how-to-scrape-dynamic-data-from-a-web-page/69164/1 "2022-10-17T18:24:44Z")

</div>

This question comes up regularly in the forum.

It usually starts with someone asking "how do I get XYZ data from this web page?". Where it turns out that the page is creating the required data dynamically using JavaScript.

There are two ways to resolve this issue. Which you use depends on how the page is working. So the first option is probably only really open to you if you can make sense of the code in your browser (view that by right-clicking on the page and selecting "Inspect" or by opening the browser's developer tools and going to the "Elements" tab.

1. Find an API call on the page that returns the data you really want
2. Use a "headless browser" to run the pages JavaScript and scrape the resulting data.

* * *

## 1) Find an API call on the page that returns the data you really want

To do this, you have to search through the web page's code or perhaps look at the network tab in the browser dev tools. Then once found, you need to see if you can call the API (which is another web endpoint) or whether it has other security that might be too hard to easily work out.

The advantage of this is that you are likely to get exactly the data you want in a form easily consumed in Node-RED (e.g. JSON or XML) and you won't have to pull the data out of some HTML.

## 2) Use a "headless browser" to run the pages JavaScript and scrape the resulting data.

There are a number of existing nodes that you may be able to use. Search for nodes that use [nbrowser](https://flows.nodered.org/search?term=nbrowser) or [puppeteer](https://flows.nodered.org/search?term=puppeteer) for example.

You could even run your own browser "headless" if you are running Node-RED on a server that has a Chromium or Firefox based browser installed. On Windows, for example, the following command line will grab a processed page using the Microsoft Edge Chromium-based browser:

```auto
"C:\Program Files (x86)\Microsoft\Edge\Application\msedge.exe" --headless --disable-gpu --enable-logging --dump-dom https://nodered.org

```

---

<div class="post-metadata">

**Author:** ![Jaxom\_99](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/jaxom_99/32/70439_2.png) [@Jaxom\_99](https://discourse.nodered.org/u/Jaxom_99)\
**Post date:** [9 November 2022 09:33 UTC](https://discourse.nodered.org/t/http-request-node-does-not-get-all-the-page-data-aka-how-to-scrape-dynamic-data-from-a-web-page/69164/2 "2022-11-09T09:33:15Z")

</div>

Thank you @TotallyInformation for this concise and clear explanation. It fits exactly what I was looking for in this forum. Could you point out to more ressources for using headless browsers ?

I have [this website of a solar production graph](https://suivideproduction.edfenr.com/customPage/efe432bd-4873-471a-8fd5) where the "hidden" API is hard to use (JWT auth through javascript, multiple requests, etc..) so I wish to use **option 2** as you describe, but I don't know where to start... I found the data as a Json object from within my browser though, so I guess it should be doable 😺

Thanks in advance to the community for any pointers 😉  
(PS : depending on answers, I could split this to a new thread...)

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [9 November 2022 15:36 UTC](https://discourse.nodered.org/t/http-request-node-does-not-get-all-the-page-data-aka-how-to-scrape-dynamic-data-from-a-web-page/69164/3 "2022-11-09T15:36:49Z")

</div>

Start here: [Library - Node-RED (nodered.org)](https://flows.nodered.org/search?term=headless)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/1X/d073cd938eafa2e558d7c2cd59003b3ef4963033.png) [@system](https://discourse.nodered.org/u/system)\
**Post date:** [16 November 2022 18:24 UTC](https://discourse.nodered.org/t/http-request-node-does-not-get-all-the-page-data-aka-how-to-scrape-dynamic-data-from-a-web-page/69164/4 "2022-11-16T18:24:46Z")

</div>

This topic was automatically closed after 30 days. New replies are no longer allowed.
