# HTML Links (Table) from website to CSV or Excel

**URL:** <https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581>\
**Category:** General\
**Created:** [17 July 2024 19:09 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581 "2024-07-17T19:09:28Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![jenssen](https://avatars.discourse-cdn.com/v4/letter/j/f0a364/32.png) [@jenssen](https://discourse.nodered.org/u/jenssen)\
**Post date:** [17 July 2024 19:09 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/1 "2024-07-17T19:09:28Z")

</div>

Hello,

Need some help to convert the table on this site: [https://www.sounds-venlo.nl/verwachte-releases/?filter-type[]=vinyl](https://www.sounds-venlo.nl/verwachte-releases/?filter-type%5B%5D=vinyl) to a CSV or Excel file.

Tried some things with the html node, but no result to retrieve the full table in a simple way.

Thanks.

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [17 July 2024 19:23 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/2 "2024-07-17T19:23:25Z")

</div>

One issue is that, though it LOOKS like a table, it really isn't.

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/2/b/2b7a46c32cf0c1e20e3e17ccadcc92da550ba1b4.png)

It is a set of links containing some spans. And the spans are defined as `display: bock` 🤷 A bit bonkers.

Your best bet is to extract the innerHTML from the `<section>` tag and then try to process it manually.

> [@jenssen](#):
>
> retrieve the full table in a simple way

There is no simple way I don't think.

---

<div class="post-metadata">

**Author:** ![craigcurtin](https://avatars.discourse-cdn.com/v4/letter/c/94ad74/32.png) [@craigcurtin](https://discourse.nodered.org/u/craigcurtin)\
**Post date:** [17 July 2024 23:36 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/3 "2024-07-17T23:36:12Z")

</div>

Unless you really have to do it in Node Red - i would suggest having a look at Power Query and the get data function in the newer versions of excel - it is a very easy and clean process to extract something into a useable excel file

Craig

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [18 July 2024 00:17 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/4 "2024-07-18T00:17:41Z")

</div>

I don't see how using PowerQuery would be any easier.

It is a matter of walking through the HTML, converting to something more useful.

---

<div class="post-metadata">

**Author:** ![bakman2](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/bakman2/32/6207_2.png) [@bakman2](https://discourse.nodered.org/u/bakman2)\
**Post date:** [18 July 2024 05:11 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/5 "2024-07-18T05:11:35Z")

</div>

This flow will produce csv output, which can be saved to a file.

It is not that hard. First get all `ahrefs` with an html node, which produces an array, use a split node and use another html node to get all `spans`, join them back into an array and pass it through a csv node.

```auto
[{"id":"e5c6200e558a48ac","type":"inject","z":"97d5eaac17934f34","name":"","props":[{"p":"payload"},{"p":"topic","vt":"str"}],"repeat":"","crontab":"","once":false,"onceDelay":0.1,"topic":"","payload":"","payloadType":"date","x":180,"y":280,"wires":[["b596b0e309430a73"]]},{"id":"b596b0e309430a73","type":"http request","z":"97d5eaac17934f34","name":"","method":"GET","ret":"txt","paytoqs":"ignore","url":"https://www.sounds-venlo.nl/verwachte-releases/?filter-type%5B%5D=vinyl","tls":"","persist":false,"proxy":"","insecureHTTPParser":false,"authType":"","senderr":false,"headers":[],"x":350,"y":280,"wires":[["3629ebe64a3b8d18"]]},{"id":"3629ebe64a3b8d18","type":"html","z":"97d5eaac17934f34","name":"","property":"payload","outproperty":"payload","tag":"body > main > section > a","ret":"html","as":"single","x":570,"y":280,"wires":[["7d336be5d72981d9"]]},{"id":"5051aa968880e841","type":"debug","z":"97d5eaac17934f34","name":"debug 595","active":true,"tosidebar":true,"console":false,"tostatus":false,"complete":"false","statusVal":"","statusType":"auto","x":1270,"y":280,"wires":[]},{"id":"466df08a13a86333","type":"html","z":"97d5eaac17934f34","name":"","property":"payload","outproperty":"payload","tag":"span","ret":"text","as":"single","x":870,"y":280,"wires":[["9a694c49dda06ad9"]]},{"id":"7d336be5d72981d9","type":"split","z":"97d5eaac17934f34","name":"","splt":"\\n","spltType":"str","arraySplt":1,"arraySpltType":"len","stream":false,"addname":"","x":750,"y":280,"wires":[["466df08a13a86333"]]},{"id":"9a694c49dda06ad9","type":"join","z":"97d5eaac17934f34","name":"","mode":"auto","build":"object","property":"payload","propertyType":"msg","key":"topic","joiner":"\\n","joinerType":"str","accumulate":"false","timeout":"","count":"","reduceRight":false,"x":990,"y":280,"wires":[["cd45a515197e7dde"]]},{"id":"cd45a515197e7dde","type":"csv","z":"97d5eaac17934f34","name":"","sep":",","hdrin":"","hdrout":"all","multi":"one","ret":"\\n","temp":"title,recordType,released,price","skip":"0","strings":true,"include_empty_strings":"","include_null_values":"","x":1110,"y":280,"wires":[["5051aa968880e841"]]}]

```

---

<div class="post-metadata">

**Author:** ![craigcurtin](https://avatars.discourse-cdn.com/v4/letter/c/94ad74/32.png) [@craigcurtin](https://discourse.nodered.org/u/craigcurtin)\
**Post date:** [18 July 2024 05:15 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/6 "2024-07-18T05:15:35Z")

</div>

Well Excel sucked it in in one go - no work - just told PQ to retrieve the URL

 ![image](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/3X/d/2/d21aba94d1da4ce73f359eda5eeb40027fce894b.png)

SO if the ultimate result is to get it to CSV looks like a pretty quick way to do so

Craig

---

<div class="post-metadata">

**Author:** ![bakman2](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/bakman2/32/6207_2.png) [@bakman2](https://discourse.nodered.org/u/bakman2)\
**Post date:** [18 July 2024 05:19 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/7 "2024-07-18T05:19:41Z")

</div>

> just told PQ to retrieve the URL

yeah excel / powerbi with powerquery works wonders, quite a strong/powerful parser

---

<div class="post-metadata">

**Author:** ![craigcurtin](https://avatars.discourse-cdn.com/v4/letter/c/94ad74/32.png) [@craigcurtin](https://discourse.nodered.org/u/craigcurtin)\
**Post date:** [18 July 2024 05:21 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/8 "2024-07-18T05:21:14Z")

</div>

Yeah - my wife (accountant) is particularly liking the ability to parse PDFs straight in

Craig

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [18 July 2024 09:57 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/9 "2024-07-18T09:57:08Z")

</div>

It's great when it works! I use PowerQuery a LOT at work to deal with large data CSV's. Didn't think to try sucking that page direct into Excel - good call.

---

<div class="post-metadata">

**Author:** ![jenssen](https://avatars.discourse-cdn.com/v4/letter/j/f0a364/32.png) [@jenssen](https://discourse.nodered.org/u/jenssen)\
**Post date:** [18 July 2024 18:26 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/10 "2024-07-18T18:26:27Z")

</div>

Really great, works like a charm!

---

<div class="post-metadata">

**Author:** ![jenssen](https://avatars.discourse-cdn.com/v4/letter/j/f0a364/32.png) [@jenssen](https://discourse.nodered.org/u/jenssen)\
**Post date:** [18 July 2024 18:28 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/11 "2024-07-18T18:28:06Z")

</div>

Is indeed simple and works too, but since I want to schedule this import/export on weekly basis, I will go for the bakman2 solution. However, this tip regarding the Excel option can be very handy in the future.

---

<div class="post-metadata">

**Author:** ![craigcurtin](https://avatars.discourse-cdn.com/v4/letter/c/94ad74/32.png) [@craigcurtin](https://discourse.nodered.org/u/craigcurtin)\
**Post date:** [19 July 2024 02:54 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/12 "2024-07-19T02:54:35Z")

</div>

Although you have a working solution - if you look at Power Bi and Power Automate (both for Windows) you can schedule tasks (such as opening excel workbooks and updating them)

Craig

---

<div class="post-metadata">

**Author:** ![TotallyInformation](https://sea2.discourse-cdn.com/flex026/user_avatar/discourse.nodered.org/totallyinformation/32/31_2.png) [@TotallyInformation](https://discourse.nodered.org/u/TotallyInformation)\
**Post date:** [21 July 2024 12:06 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/13 "2024-07-21T12:06:11Z")

</div>

> [@craigcurtin](#):
>
> if you look at Power Bi and Power Automate (both for Windows) you can schedule tasks (such as opening excel workbooks and updating them)

Assuming you have appropriate licensing of course. ☹

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex026/uploads/nodered/original/1X/d073cd938eafa2e558d7c2cd59003b3ef4963033.png) [@system](https://discourse.nodered.org/u/system)\
**Post date:** [4 August 2024 12:06 UTC](https://discourse.nodered.org/t/html-links-table-from-website-to-csv-or-excel/89581/14 "2024-08-04T12:06:13Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
