# Web scraping to grab an items price, but coming back empty

**URL:** <https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894>\
**Category:** General\
**Created:** [27 October 2020 14:26 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894 "2020-10-27T14:26:19Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![clandestine-avocado](https://avatars.discourse-cdn.com/v4/letter/c/5fc32e/32.png) [@clandestine-avocado](https://discourse.nodered.org/u/clandestine-avocado)\
**Post date:** [27 October 2020 14:26 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/1 "2020-10-27T14:26:19Z")

</div>

I'm trying to grab price data from [this page](https://www.lowes.com/pd/Common-6-in-x-6-in-x-20-ft-Actual-5-5-in-x-5-5-in-x-20-ft-2-Treated-Lumber/1000698158) via HTTP request. The price appears to be in `<div class="sc-fznMnq biVOqy">$146.63 </div>`. So, in the HTML Node, I set the _Selector_ field to .sc-fznMnq biVOqy (following the [example from the docs](https://cookbook.nodered.org/http/simple-get-request)). But it keeps coming back empty - am I missing something here?

Here is my flow:

`[{"id":"e5ce3a5d.b3d348","type":"inject","z":"f20f1b13.5e8608","name":"make request","repeat":"","crontab":"","once":false,"topic":"","payload":"","payloadType":"date","x":1310,"y":300,"wires":[["251988f5.6a3478"]]},{"id":"251988f5.6a3478","type":"http request","z":"f20f1b13.5e8608","name":"6x6x20","method":"GET","ret":"txt","paytoqs":"ignore","url":"https://www.lowes.com/pd/Common-6-in-x-6-in-x-20-ft-Actual-5-5-in-x-5-5-in-x-20-ft-2-Treated-Lumber/1000698158","tls":"","persist":false,"proxy":"","authType":"","x":1464.5,"y":300,"wires":[["15cd391f.bee797","ae2bd9d0.9ba898"]]},{"id":"bfbea6ff.1329b8","type":"debug","z":"f20f1b13.5e8608","name":"","active":true,"console":"false","complete":"false","x":1970,"y":300,"wires":[]},{"id":"15cd391f.bee797","type":"html","z":"f20f1b13.5e8608","name":"","property":"","outproperty":"","tag":".sc-fznMnq biVOqy","ret":"text","as":"single","x":1710,"y":300,"wires":[["bfbea6ff.1329b8"]]}]`

---

<div class="post-metadata">

**Author:** ![Steve-Mcl](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/steve-mcl/32/4826_2.png) [@Steve-Mcl](https://discourse.nodered.org/u/Steve-Mcl)\
**Post date:** [27 October 2020 14:37 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/2 "2020-10-27T14:37:41Z")

</div>

It is likely the value is populated separate from the html (i.e. via JavaScript). In otherwords, the value is possibly not present in the html until some time after is arrives in the browser.

You might need to look at the network tab in dev tools to see if you can find the source of the data.

---

<div class="post-metadata">

**Author:** ![Nodi.Rubrum](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/nodi.rubrum/32/107482_2.png) [@Nodi.Rubrum](https://discourse.nodered.org/u/Nodi.Rubrum)\
**Post date:** [27 October 2020 18:35 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/3 "2020-10-27T18:35:21Z")

</div>

If the web page has frames... make sure you are getting the right frame. I wrote a cable modem status scraper in NR, and I was not getting the right results, realized looking at the HTML details, the status value was in a nested frame, which was not obvious. I just set the HTTP request to the right frame, and bingo, got the cable modem status as expected.

---

<div class="post-metadata">

**Author:** ![Steve-Mcl](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/steve-mcl/32/4826_2.png) [@Steve-Mcl](https://discourse.nodered.org/u/Steve-Mcl)\
**Post date:** [27 October 2020 18:50 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/4 "2020-10-27T18:50:20Z")

</div>

As i thought, value is NOT present in the HTML (it is updated in a later background call to another URL)

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/e/2/e2dc084ad2de81def37abe7bd87cece48d4e5e96.png)

The actual URL is `https://www.lowes.com/pd/1000698158/productdetail/2209/Guest` so do a request to that URL, set the request node to JSON then access the values from `msg.payload.productDetails["1000698158"].price.itemPrice`

---

<div class="post-metadata">

**Author:** ![clandestine-avocado](https://avatars.discourse-cdn.com/v4/letter/c/5fc32e/32.png) [@clandestine-avocado](https://discourse.nodered.org/u/clandestine-avocado)\
**Post date:** [27 October 2020 21:08 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/5 "2020-10-27T21:08:08Z")

</div>

@Steve-Mcl great - I can see the price value if I throw that URL in a browser - but price comes back as null in the JSON response via NR. Oddly, all the other data (that I don't care about!) is in the JSON response.

Via Browser:  
 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/5/d/5d1606b281032fca600ea66e7b4f69a378ba0d61.png)

Via NR:

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/b/c/bc0a86c9758d6e1ec4cbeb461de4badd8cc42a78.png)

---

<div class="post-metadata">

**Author:** ![Steve-Mcl](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/steve-mcl/32/4826_2.png) [@Steve-Mcl](https://discourse.nodered.org/u/Steve-Mcl)\
**Post date:** [27 October 2020 21:38 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/6 "2020-10-27T21:38:58Z")

</div>

At a guess, you need to replicate the headers Chrome uses.

Or set cookies.

---

<div class="post-metadata">

**Author:** ![clandestine-avocado](https://avatars.discourse-cdn.com/v4/letter/c/5fc32e/32.png) [@clandestine-avocado](https://discourse.nodered.org/u/clandestine-avocado)\
**Post date:** [28 October 2020 00:56 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/7 "2020-10-28T00:56:09Z")

</div>

Thanks @Steve-Mcl I'll look into that. Not something I am familiar with - so if you have any good resources to point me in the right direction, I'd be grateful. If not, thanks for the assistance so far - I'm closer than I was this morning!

---

<div class="post-metadata">

**Author:** ![Steve-Mcl](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/steve-mcl/32/4826_2.png) [@Steve-Mcl](https://discourse.nodered.org/u/Steve-Mcl)\
**Post date:** [28 October 2020 07:21 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/8 "2020-10-28T07:21:56Z")

</div>

Make a request to the original page - you will get cookie(s) then make 2nd request to the JSON URL with the cookies from the 1st response in the 2nd request.

Basically check using Devtools what is sent in the request for the JSON & try to replicate that in node-red.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/1X/d073cd938eafa2e558d7c2cd59003b3ef4963033.png) [@system](https://discourse.nodered.org/u/system)\
**Post date:** [27 December 2020 07:22 UTC](https://discourse.nodered.org/t/web-scraping-to-grab-an-items-price-but-coming-back-empty/34894/9 "2020-12-27T07:22:12Z")

</div>

This topic was automatically closed 60 days after the last reply. New replies are no longer allowed.
