# Web scraping text and looking result up in table

**URL:** <https://discourse.nodered.org/t/web-scraping-text-and-looking-result-up-in-table/64743>\
**Category:** General\
**Tags:** http-request\
**Created:** [5 July 2022 19:49 UTC](https://discourse.nodered.org/t/web-scraping-text-and-looking-result-up-in-table/64743 "2022-07-05T19:49:00Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![EspeeFunsail](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/espeefunsail/32/64244_2.png) [@EspeeFunsail](https://discourse.nodered.org/u/EspeeFunsail)\
**Post date:** [5 July 2022 19:49 UTC](https://discourse.nodered.org/t/web-scraping-text-and-looking-result-up-in-table/64743/1 "2022-07-05T19:49:00Z")

</div>

I am using http-request to get the text from this website:  
[https://ewn.co.za/assets/loadshedding/api/status](https://ewn.co.za/assets/loadshedding/api/status)

All I want to do is get the data after a certain set of words, for example: "City Customers: Stage " and then whatever number is located there.

From there I want to lookup the number in this table:

> **[Load\_Shedding\_All\_Areas\_Schedule\_and\_Map.pdf](https://resource.capetown.gov.za/documentcentre/Documents/Procedures,%20guidelines%20and%20regulations/Load_Shedding_All_Areas_Schedule_and_Map.pdf)**
>
> 1935.87 KB

and get time slots which shows when my area will be load shed (Welcome to Africa).

Any assistance would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![UnborN](https://avatars.discourse-cdn.com/v4/letter/u/4491bb/32.png) [@UnborN](https://discourse.nodered.org/u/UnborN)\
**Post date:** [5 July 2022 22:34 UTC](https://discourse.nodered.org/t/web-scraping-text-and-looking-result-up-in-table/64743/2 "2022-07-05T22:34:01Z")

</div>

> [@EspeeFunsail](#):
>
> for example: "City Customers: Stage " and then whatever number is located there.

You could use [regular expressions](https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/String/match#try_it) using a capture group in a Function node  
to filter out the number from the text.

```auto
msg.payload = msg.payload.match(/CITY CUSTOMERS: STAGE (\d)+ /)
msg.payload = Number(msg.payload[1]) // get 1st regex group and convert to number

return msg;

```

* * *

**Test Flow:**

```auto
[{"id":"87b731dda38031a3","type":"inject","z":"54efb553244c241f","name":"","props":[{"p":"payload"},{"p":"topic","vt":"str"}],"repeat":"","crontab":"","once":false,"onceDelay":0.1,"topic":"","payload":"","payloadType":"date","x":280,"y":1520,"wires":[["81b689f8a88f21dc"]]},{"id":"81b689f8a88f21dc","type":"http request","z":"54efb553244c241f","name":"","method":"GET","ret":"txt","paytoqs":"ignore","url":"https://ewn.co.za/assets/loadshedding/api/status","tls":"","persist":false,"proxy":"","authType":"","senderr":false,"headers":[],"x":450,"y":1520,"wires":[["ca748687180a5ae1","23dd45c26e3acc66"]]},{"id":"ca748687180a5ae1","type":"function","z":"54efb553244c241f","name":"function 1","func":"\nmsg.payload = msg.payload.match(/CITY CUSTOMERS: STAGE (\\d)+ /)\nmsg.payload = Number(msg.payload[1]) // get 1st regex group and convert to number\n\nreturn msg;","outputs":1,"noerr":0,"initialize":"","finalize":"","libs":[],"x":620,"y":1520,"wires":[["8bfb25d4f46c7fef"]]},{"id":"23dd45c26e3acc66","type":"debug","z":"54efb553244c241f","name":"debug 1","active":true,"tosidebar":true,"console":false,"tostatus":false,"complete":"false","statusVal":"","statusType":"auto","x":620,"y":1460,"wires":[]},{"id":"8bfb25d4f46c7fef","type":"debug","z":"54efb553244c241f","name":"debug 2","active":true,"tosidebar":true,"console":false,"tostatus":false,"complete":"false","statusVal":"","statusType":"auto","x":800,"y":1520,"wires":[]}]

```

regarding matching that number from a pdf file .. for that i have no idea  
maybe copy paste and patiently edit the data to a JS array that you can later use as a lookup table.

---

<div class="post-metadata">

**Author:** ![EspeeFunsail](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/espeefunsail/32/64244_2.png) [@EspeeFunsail](https://discourse.nodered.org/u/EspeeFunsail)\
**Post date:** [6 July 2022 18:57 UTC](https://discourse.nodered.org/t/web-scraping-text-and-looking-result-up-in-table/64743/3 "2022-07-06T18:57:05Z")

</div>

Thank you. The regex expression worked. First time using it as I'm very new to Node Red.

Next up is to figure out the JS array and looking up the stage and time slots from it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/1X/d073cd938eafa2e558d7c2cd59003b3ef4963033.png) [@system](https://discourse.nodered.org/u/system)\
**Post date:** [4 September 2022 18:57 UTC](https://discourse.nodered.org/t/web-scraping-text-and-looking-result-up-in-table/64743/4 "2022-09-04T18:57:09Z")

</div>

This topic was automatically closed 60 days after the last reply. New replies are no longer allowed.
