# HTML node parser ; how to use the selector function to lookup a value?

**URL:** <https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953>\
**Category:** General\
**Created:** [8 October 2020 16:14 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953 "2020-10-08T16:14:20Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![ChillXXL](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/chillxxl/32/30408_2.png) [@ChillXXL](https://discourse.nodered.org/u/ChillXXL)\
**Post date:** [8 October 2020 16:14 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/1 "2020-10-08T16:14:21Z")

</div>

I try to extract a value from a website and use the HTTP (get) node connected to a HTML node.  
I cannot figure out how to extract a value on the page.

The value I want to retrieve is "F56386285397M4Y403" and is on the buttom of the page which will be different each time te page is refreshed.

On the bottom of the HTML is the value located:

```auto
*<....*
*<script type="text/javascript">*
*// <![CDATA[*
*jQuery(document).ready(function() {liftAjax.lift_successRegisterGC();});*
*var lift_page = "F56386285397M4Y403";*
*// ]]>*
*</script></body>*
*</html>*

```

I know the path is :"/html/body/script[10]/text()"

How can I point the HTML node to this path and retrieve the value of var "lift page"?  
(I do receive input in the HTML node from the HTTP node)

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [8 October 2020 17:02 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/2 "2020-10-08T17:02:02Z")

</div>

The problem is that your value is not in the DOM so you cannot use a simple selector to get hold of it.

The simples approach is to read the html page as text and then apply a change node using a regular expression to grab the text.

This expression for example, will find the text and store it in a token:

```auto
/\*var lift_page = "(.*)"/

```

You can use a replace option in a change node to just output `$1` which will return the code on its own.

---

<div class="post-metadata">

**Author:** ![ChillXXL](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/chillxxl/32/30408_2.png) [@ChillXXL](https://discourse.nodered.org/u/ChillXXL)\
**Post date:** [9 October 2020 08:06 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/3 "2020-10-09T08:06:47Z")

</div>

Thanks! Now I understand why it didn't worked because of the text is not in the DOM.

I cannot make the change node to work. The text is output by the HTML node (a string with132 characters; just like the light grey text in above post):

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/d/3/d3b19c8d2d2e1c90e9f7c3fa6dbcc933c7027e87.png)

and the change node connected to the HTML node is adjusted like:

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/6/b/6ba869e8bf63891263e39b6016dd38dd9b518f6c.png)

The change node will not replace but outputs the complete msg. Can you see what is wrong?  
(PS can a split in a function also work?)

---

<div class="post-metadata">

**Author:** ![zenofmud](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/zenofmud/32/316_2.png) [@zenofmud](https://discourse.nodered.org/u/zenofmud)\
**Post date:** [9 October 2020 08:15 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/4 "2020-10-09T08:15:52Z")

</div>

In your first post you have

```auto
*<....*
*<script type="text/javascript">*
*// <![CDATA[*
*jQuery(document).ready(function() {liftAjax.lift_successRegisterGC();});*
*var lift_page = "F56386285397M4Y403";*
*// ]]>*
*</script></body>*
*</html>*

```

does each line actually have an asterisk as the first character? If ther is no asterick then you need to change your change node to reflect that.

---

<div class="post-metadata">

**Author:** ![ChillXXL](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/chillxxl/32/30408_2.png) [@ChillXXL](https://discourse.nodered.org/u/ChillXXL)\
**Post date:** [9 October 2020 08:19 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/5 "2020-10-09T08:19:52Z")

</div>

The "\*" are put in by this NodeRed forum post app. with the "\</\>" button. The actual output of the HTML node is now:

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/b/4/b41fb40117e764d2ee3f8414291a8822f7425393.png)

---

<div class="post-metadata">

**Author:** ![zenofmud](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/zenofmud/32/316_2.png) [@zenofmud](https://discourse.nodered.org/u/zenofmud)\
**Post date:** [9 October 2020 09:00 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/6 "2020-10-09T09:00:48Z")

</div>

> [@ChillXXL](#):
>
> The "\*" are put in by this NodeRed forum post app. with the "\</\>" button.

Hmmm not when I do it

```auto
<....*
<script type="text/javascript">*
// <![CDATA[*
jQuery(document).ready(function() {liftAjax.lift_successRegisterGC();});*
var lift_page = "F56386285397M4Y403";*
// ]]>*
</script></body>*
</html>*

```

anyways, in your regex you have `/\*var lift_page = "(.*)"/` not being a regex expert (it makes my brain hurt) should you have the begining`\*`?

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [9 October 2020 19:11 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/7 "2020-10-09T19:11:05Z")

</div>

> [@zenofmud](#):
>
> should you have the begining `\*` ?

Nope! And regex hurts less than JSONata 😃 Anyway there are lots of useful online regex testers.

The following is enough anyway:

```auto
/lift_page = "(.*)"/

```

And I wouldn't bother with the html extract, just send it the whole page text.

---

<div class="post-metadata">

**Author:** ![ChillXXL](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/chillxxl/32/30408_2.png) [@ChillXXL](https://discourse.nodered.org/u/ChillXXL)\
**Post date:** [10 October 2020 21:13 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/8 "2020-10-10T21:13:42Z")

</div>

> [@ChillXXL](#):
>
> The change node will not replace but outputs the complete msg. Can you see what is wrong?  
> (PS can a split in a function also work?)

still the whole ouput. Is this setting ok (see highlight)?:  
 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/4/9/491c25cf69d012c7cbd256800a51f9cff963417f.png)

**update: replace with [$] "env variable" is not the right way; should be a string.**

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [11 October 2020 00:02 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/9 "2020-10-11T00:02:18Z")

</div>

You need:

```auto
.*lift_page = "(.*)".*

```

I forgot that you don't need the regex initial/trailing slashes and you are trying to _replace_ everything except the value so you need to include everything before and after as well.

---

<div class="post-metadata">

**Author:** ![ChillXXL](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/chillxxl/32/30408_2.png) [@ChillXXL](https://discourse.nodered.org/u/ChillXXL)\
**Post date:** [11 October 2020 13:44 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/10 "2020-10-11T13:44:44Z")

</div>

> [@TotallyInformation](#):
>
> `.*lift_page = "(.*)".*`

Thanks for the adjustment. It is a bit better but not yet perfect. It replaces the whole 'var lift page....' bit for only the code I need. But how can I het rid off all the other parts before and after. The output is now:

```auto
// <![CDATA[
jQuery(document).ready(function() {liftAjax.lift_successRegisterGC();});
F6057066220042K3GSG
// ]]>

```

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [11 October 2020 15:32 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/11 "2020-10-11T15:32:24Z")

</div>

Not in my test it isn't and I don't know how it could be since `.*lift_page = "` selects everything before the value and all of that is thrown away when the replace value is just $1

The main things that might be improved in that regex are:

1. It is possible that `"(.*)"` might actually select too much since regex is "geedy" by default. Won't happen with the example text you've shown though.
2. Selecting for `"` might fail if the author of the page decides to switch to default single quotes instead of double.

Best thing to do is to find a regex testing website, paste the HTML source into it and try out the regex.

---

<div class="post-metadata">

**Author:** ![ChillXXL](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/chillxxl/32/30408_2.png) [@ChillXXL](https://discourse.nodered.org/u/ChillXXL)\
**Post date:** [11 October 2020 17:11 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/12 "2020-10-11T17:11:40Z")

</div>

Thanks again. I think that I know what is going on why it isn't working. The tekst I quoted before, see this post below, was by clicking on the node and copy the string. BUT when I just look at the node, see picture below, little "Carriage Return" symbols are show (see yellow highlite):  
 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/1/3/134502eeea4a9a2b2e1e0ea8fddab19cc2f828c1.png)

It seems that these carriage returns somehow break the replacement node. I say this because the word "var" in front of "var lift\_page" is removed correctly with your supplied regex formula.  
See above output of the change node. What do you think?

I can not change the output of the HTML node any better than this:

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/5/9/594f35e3136639e379c4e7ffc4b095ce5c9e1761.png)

> [@ChillXXL](#):
>
> ```auto
> // <![CDATA[
> jQuery(document).ready(function() {liftAjax.lift_successRegisterGC();});
> F6057066220042K3GSG
> // ]]>
> 
> ```

In the meanwhile I'll used a function node with a split and that works correctly:

```auto
msg.payload = msg.payload.split(" ")[5].substr(1,18);
flow.set("endpointid", msg.payload) //set endpointid
return msg;

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/1X/d073cd938eafa2e558d7c2cd59003b3ef4963033.png) [@system](https://discourse.nodered.org/u/system)\
**Post date:** [10 December 2020 17:11 UTC](https://discourse.nodered.org/t/html-node-parser-how-to-use-the-selector-function-to-lookup-a-value/33953/13 "2020-12-10T17:11:40Z")

</div>

This topic was automatically closed 60 days after the last reply. New replies are no longer allowed.
