Getting a web table into Excel: the four things that go wrong
Published by Quillfold, which makes TableHarvest.
Most people get a web table into a spreadsheet in one of four ways: select and copy it, use Excel’s Data > From Web, use Google Sheets’ IMPORTHTML, or use a browser extension. Each works well on simple tables. On the others, the same four problems come up again and again.
Each guide below explains why the problem happens, how to check for it in the browser, how to fix it by hand, what Excel and Google Sheets can do, and what extensions do. Where another method is enough, the guide says so.
The four problems
Rows go missing
The page says 57 entries and you have 10, or a long list stops partway. The usual causes: the rest is on other pages or behind Load more, the grid only draws the rows on screen, or the rows arrive after the tool has read the page. Start by finding the total the page states, so you have something to compare with.
The next page isn’t captured
Your export has page 1 only, or page 1 twice. “Next” can be a link, a button that swaps the table, a script that reloads the whole page, or an arrow with no text. Each kind breaks a different tool. Hovering over Next and watching the address bar tells you which kind you have.
CSV opens garbled, or numbers change
Accented or Asian text turns into symbols, 00123 becomes 123, a 20-digit ID becomes 1.23E+19. The first is an encoding problem with a two-minute fix; the second is Excel converting values, which Microsoft 365 lets you switch off. Exporting XLSX instead of CSV avoids both.
Headers come out as “Column 1”, or merged cells break
The header row isn’t recognised, merged cells leave blanks, a two-row header becomes data, or values shift into the next column. These come from how the table is written in HTML. Power Query’s Fill Down and Use First Row as Headers fix the common cases after import.
The built-in tools
- Excel “From Web” can’t see your table: the steps when it works, which Excel versions have it (Windows does; Mac and Excel for the web don’t list it), and what Microsoft says about pages built by scripts, sign-in and paging.
- IMPORTHTML returns an error or the wrong table: how the index is counted, the errors Google documents, and the hourly refresh.
Which method fits (as of October 2026)
| Situation | Copy and paste | Excel From Web (Windows) | Google Sheets IMPORTHTML | Browser extension |
|---|---|---|---|---|
| One page, table in the page source | Yes | Yes | Yes | Yes |
| Must refresh later | No | Yes (refresh the query) | Yes (hourly while open) | Varies; TableHarvest doesn’t |
Pages with their own addresses (?page=2) | One page at a time | One query per page, then Append | One formula per page | Instant Data Scraper: listed. Table Capture: paid tier. TableHarvest: up to 3 pages per capture free, more with Pro |
| Next, Load more or infinite scroll without new addresses | One page at a time | No | No | Same as the row above |
| Table built by scripts after loading | Yes, what’s on screen | Microsoft warns results can be inconsistent; doesn’t say whether Excel runs scripts | Not documented by Google | Yes, they read the open page |
| Page you must sign in to | Yes | Web.Page lists Anonymous, Windows, Web API only | Not documented by Google | Yes, they read the open page |
| Mac or Excel for the web | Yes | Not listed | Yes | Yes (Chrome) |
Details and sources are in each guide. “Not documented” means the vendor’s help pages we read on 7 October 2026 don’t say either way.
If the site has its own download
Check for an Export, Download or CSV button near the table, or an “open data” page on the site, before anything else. Statistics offices and dashboards sometimes offer one. It is the most complete source, and you don’t need any of the methods above.
Sources
The sources for each statement are listed at the end of each guide.