# Question about extraction of text on HTML webpage

**URL:** <https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321>\
**Category:** General\
**Created:** [20 October 2022 09:58 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321 "2022-10-20T09:58:07Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![polarrys](https://avatars.discourse-cdn.com/v4/letter/p/a3d4f5/32.png) [@polarrys](https://discourse.nodered.org/u/polarrys)\
**Post date:** [20 October 2022 09:58 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/1 "2022-10-20T09:58:07Z")

</div>

Dear Node-red community.

I read a lot of post on how to use html request and html node and I can't do that I want. Let me explain :  
I want to extract and to post 2 text on my dashbord from this web site :

[RTE Tempo](https://www.services-rte.com/fr/visualisez-les-donnees-publiees-par-rte/calendrier-des-offres-de-fourniture-de-type-tempo.html)

The two value is in red on my picture

 ![Sans titre](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/0/7/0752db21caf6f29c14ecf8ae2adc90f9dd121005.jpeg)

I save the html webpage on my computer and I see the two value :

 ![Sans titre2](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/6/4/64ddcac3584f7018b2f0e255fd71069af8464051.jpeg)

Now, I want to extract this 2 text value with node-red. I test a lot of combinaison but It doesn't work.

Can you help me on this problem.

Thanks a lot.

---

<div class="post-metadata">

**Author:** ![janvda](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/janvda/32/234_2.png) [@janvda](https://discourse.nodered.org/u/janvda)\
**Post date:** [20 October 2022 14:08 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/2 "2022-10-20T14:08:58Z")

</div>

1. First you need to get the page source (html code) in node-red.  
This you can do with the `http request` node.

2. Then you need to extract the relevant info from the page source.  
You can do this by using an `html` node and specifying the appropriate css-selector.

The appropriate css-selector can be found by opening the page in chrome.  
Select the field `Jour blue`, right click and select `inspect`.  
Then in the elements tab it selects the corresponding page source.  
If you then select that line in the elements tab and right click and select `Copy` \> `Copy Selector"` It will copy the css selector to your clipboard which you then can paste into the configuration of your node-red `html` node.

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/4/9/49fd00ce759fabc7b29ac73f7cfece318546e3e6.png)

Using the same chrome developer tools window you can also test if a css selector is selecting the appropriate information.  
For that press Command + Option + F (Mac) or Control + Shift + F (Windows/Linux) in the elements tab to open the search bar at the bottom.  
In that search bar you can then paste the css selector and validate if it is selecting the correct thing.

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/2/d/2df7129266fdc85d6af0e47a4eb7cd6dbe457d44.png)

---

<div class="post-metadata">

**Author:** ![polarrys](https://avatars.discourse-cdn.com/v4/letter/p/a3d4f5/32.png) [@polarrys](https://discourse.nodered.org/u/polarrys)\
**Post date:** [20 October 2022 17:15 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/3 "2022-10-20T17:15:04Z")

</div>

Dear janvda,

Thanks you very much for your help.  
I use Chrome to copy the CSS code and I include it in the html note. But, my msg.payload say : [Empty]

There is my export node :

```auto
[
    {
        "id": "74c8759c107a9b82",
        "type": "inject",
        "z": "3ae78e8265bcb79e",
        "name": "make request",
        "repeat": "",
        "crontab": "",
        "once": false,
        "topic": "",
        "payload": "",
        "payloadType": "date",
        "x": 150,
        "y": 1320,
        "wires": [
            [
                "da6e7a307965c4f1"
            ]
        ]
    },
    {
        "id": "da6e7a307965c4f1",
        "type": "http request",
        "z": "3ae78e8265bcb79e",
        "name": "",
        "method": "GET",
        "ret": "txt",
        "paytoqs": "ignore",
        "url": "https://www.services-rte.com/fr/visualisez-les-donnees-publiees-par-rte/calendrier-des-offres-de-fourniture-de-type-tempo.html",
        "tls": "",
        "persist": false,
        "proxy": "",
        "insecureHTTPParser": false,
        "authType": "",
        "senderr": false,
        "headers": [],
        "x": 310,
        "y": 1440,
        "wires": [
            [
                "78a76b655807cd90"
            ]
        ]
    },
    {
        "id": "a3878f8e23a7bb9f",
        "type": "debug",
        "z": "3ae78e8265bcb79e",
        "name": "",
        "active": true,
        "tosidebar": true,
        "console": false,
        "tostatus": false,
        "complete": "payload",
        "targetType": "msg",
        "statusVal": "",
        "statusType": "auto",
        "x": 1110,
        "y": 1440,
        "wires": []
    },
    {
        "id": "78a76b655807cd90",
        "type": "html",
        "z": "3ae78e8265bcb79e",
        "name": "",
        "property": "",
        "outproperty": "",
        "tag": "#wrapper > div > div > div.c-editorial-page __container > div.c-editorial-page__ content > tempo > div > div.o-grid.c-tempo > div:nth-child(1) > div.c-tempo __bloc-1__ body.c-tempo__background--blue",
        "ret": "text",
        "as": "single",
        "x": 780,
        "y": 1640,
        "wires": [
            [
                "a3878f8e23a7bb9f"
            ]
        ]
    }
]

```

What is worng ?

Thanks

---

<div class="post-metadata">

**Author:** ![janvda](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/janvda/32/234_2.png) [@janvda](https://discourse.nodered.org/u/janvda)\
**Post date:** [20 October 2022 20:53 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/4 "2022-10-20T20:53:55Z")

</div>

> [@polarrys](#):
>
> What is worng ?

The page source retrieved by the http request node is not the same as the page source you see in your browser.  
This is because for that page your browser will do several other GET requests to construct your page.

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/b/b/bb21a265169504f8b6c1530b4bb4c46a07447f33.jpeg)

I found following request in the list of GET requests:

- [https://www.services-rte.com/cms/open\_data/v1/tempo?season=2022-2023](https://www.services-rte.com/cms/open_data/v1/tempo?season=2022-2023)

I am wondering if that request is not directly returning the information you are looking for.

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/7/e/7ea27044d96709dfd4bb64b5e6c55e3bbe5ad6f6.png)

---

<div class="post-metadata">

**Author:** ![polarrys](https://avatars.discourse-cdn.com/v4/letter/p/a3d4f5/32.png) [@polarrys](https://discourse.nodered.org/u/polarrys)\
**Post date:** [21 October 2022 07:44 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/5 "2022-10-21T07:44:23Z")

</div>

Dear janvda,

Thanks for your message.  
Actually, I want to extract the color day (like blue) and the tomorrow color (like blue).  
If I check your link, I have only a calendar view.

It become very complicated for me lol. I'm a newbie with node-red. Maybe, it exist another way to simply extract this 2 values ?

---

<div class="post-metadata">

**Author:** ![janvda](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/janvda/32/234_2.png) [@janvda](https://discourse.nodered.org/u/janvda)\
**Post date:** [21 October 2022 11:37 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/6 "2022-10-21T11:37:36Z")

</div>

> [@polarrys](#):
>
> Actually, I want to extract the color day (like blue) and the tomorrow color (like blue).  
> If I check your link, I have only a calendar view.

Indeed the link gives information for other dates as well, but I think that the link also contains the information you want for today and tomorrow.

Here below a flow that is extracting the information for today and tomorrow from url:

- [https://www.services-rte.com/cms/open\_data/v1/tempo?season=2022-2023](https://www.services-rte.com/cms/open_data/v1/tempo?season=2022-2023)

```auto
[
    {
        "id": "c1771c42fd752d19",
        "type": "inject",
        "z": "a30b805e19d6ab43",
        "name": "",
        "props": [
            {
                "p": "payload"
            },
            {
                "p": "topic",
                "vt": "str"
            }
        ],
        "repeat": "",
        "crontab": "",
        "once": false,
        "onceDelay": 0.1,
        "topic": "",
        "payload": "",
        "payloadType": "date",
        "x": 190,
        "y": 160,
        "wires": [
            [
                "0c424a93e087991e"
            ]
        ]
    },
    {
        "id": "0c424a93e087991e",
        "type": "http request",
        "z": "a30b805e19d6ab43",
        "name": "",
        "method": "GET",
        "ret": "txt",
        "paytoqs": "ignore",
        "url": "https://www.services-rte.com/cms/open_data/v1/tempo?season=2022-2023",
        "tls": "",
        "persist": false,
        "proxy": "",
        "insecureHTTPParser": false,
        "authType": "",
        "senderr": false,
        "headers": [],
        "x": 370,
        "y": 160,
        "wires": [
            [
                "c5b0074a60e66be6"
            ]
        ]
    },
    {
        "id": "558948738112c46b",
        "type": "debug",
        "z": "a30b805e19d6ab43",
        "name": "json output",
        "active": true,
        "tosidebar": true,
        "console": false,
        "tostatus": false,
        "complete": "payload",
        "targetType": "msg",
        "statusVal": "",
        "statusType": "auto",
        "x": 610,
        "y": 120,
        "wires": []
    },
    {
        "id": "c5b0074a60e66be6",
        "type": "json",
        "z": "a30b805e19d6ab43",
        "name": "",
        "property": "payload",
        "action": "",
        "pretty": false,
        "x": 550,
        "y": 160,
        "wires": [
            [
                "558948738112c46b",
                "cfa7222b2b9f5c46"
            ]
        ]
    },
    {
        "id": "cfa7222b2b9f5c46",
        "type": "change",
        "z": "a30b805e19d6ab43",
        "name": "extract today and tomorrow colors",
        "rules": [
            {
                "t": "set",
                "p": "payload",
                "pt": "msg",
                "to": "( \t $today:=$now(\"[Y0001]-[M01]-[D01]\");\t $tomorrow:=(($now()~>$toMillis())+(24*60*60*1000))~>$fromMillis(\"[Y0001]-[M01]-[D01]\");\t {\t \"today\" : $today,\t \"tomorrow\" : $tomorrow,\t \"today_color\" : payload.values~>$lookup($today),\t \"tomorrow_color\" : payload.values~>$lookup($tomorrow),\t \"today_fallback\" : payload.values~>$lookup($today & \"-fallback\"),\t \"tomorrow_fallback\" : payload.values~>$lookup($tomorrow & \"-fallback\")\t }\t)",
                "tot": "jsonata"
            }
        ],
        "action": "",
        "property": "",
        "from": "",
        "to": "",
        "reg": false,
        "x": 800,
        "y": 160,
        "wires": [
            [
                "a49b534377d9bfcb"
            ]
        ]
    },
    {
        "id": "a49b534377d9bfcb",
        "type": "debug",
        "z": "a30b805e19d6ab43",
        "name": "today / tomorrow color",
        "active": true,
        "tosidebar": true,
        "console": false,
        "tostatus": false,
        "complete": "payload",
        "targetType": "msg",
        "statusVal": "",
        "statusType": "auto",
        "x": 1000,
        "y": 120,
        "wires": []
    }
]

```

So this is the output if I trigger the above flow in the node-red editor:

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/c/6/c6bcf9ee89fb5f6ff7bbe33cabf437974f47b8c2.png)

---

<div class="post-metadata">

**Author:** ![polarrys](https://avatars.discourse-cdn.com/v4/letter/p/a3d4f5/32.png) [@polarrys](https://discourse.nodered.org/u/polarrys)\
**Post date:** [21 October 2022 12:23 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/7 "2022-10-21T12:23:32Z")

</div>

janvda, Just Waow. It work fine for me.

Now, I must arrange the text to show only the 2 color on my dashboard.

Janvda, thanks you very much for your help.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/1X/d073cd938eafa2e558d7c2cd59003b3ef4963033.png) [@system](https://discourse.nodered.org/u/system)\
**Post date:** [4 November 2022 12:23 UTC](https://discourse.nodered.org/t/question-about-extraction-of-text-on-html-webpage/69321/8 "2022-11-04T12:23:49Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
