# Tab Separated Value parser to supplement CSV?

**URL:** <https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829>\
**Category:** Help\
**Tags:** data-import\
**Created:** [September 15, 2021, 3:23pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829 "2021-09-15T15:23:20Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![emyoulation](https://yyz2.discourse-cdn.com/free1/user_avatar/gramps.discourse.group/emyoulation/32/67_2.png) [@emyoulation](https://gramps.discourse.group/u/emyoulation)\
**Post date:** [September 15, 2021, 3:23pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829/1 "2021-09-15T15:23:20Z")

</div>

I was experimenting with transferring my Place hierarchy using the [CSV view export](https://www.gramps-project.org/wiki/index.php/Gramps_5.1_Wiki_Manual_-_Settings#Export_View) & the [Import Text](https://www.gramps-project.org/wiki/index.php/ImportGramplet) Gramplet. But one common step is to clean up the data a bit in Excel.

I was _hoping_ to be able to just copy filtered chunks of spreadsheet & have it be parsed in a variant of the **Import Text Gramplet** … _without_ messing with double-quoting Place Titles & names. And _without_ applying an Excel formula to concatenate each row with comma delimiters or having to write the rows to a subsequently importable CSV file.

But the the **Import Text Gramplet** chokes on tab separated text. And our importer doesn’t recognize `.tsv` file extensions nor tab delimited content.

There’s a [thread on StackOverflow](https://stackoverflow.com/questions/42358259/how-to-parse-tsv-file-with-python/42362801#42362801) that says the python [csv module](https://docs.python.org/3/library/csv.html) can delimit on tabs instead of commas.

```python
with open("file.tsv") as fd:
    rd = csv.reader(fd, delimiter="\t", quotechar='"')
    for row in rd:
        print(row)

```

Is there a way to support a different delimiter in our text parser?

---

<div class="post-metadata">

**Author:** ![prculley](https://yyz2.discourse-cdn.com/free1/user_avatar/gramps.discourse.group/prculley/32/41_2.png) [@prculley](https://gramps.discourse.group/u/prculley)\
**Post date:** [September 17, 2021, 2:40pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829/2 "2021-09-17T14:40:46Z")

</div>

It looks like you have found such a way. Since the Import Text Gramplet is apparently an addon, why don’t you try modifying it to work the way you want, and then submit the changes as a PR?

---

<div class="post-metadata">

**Author:** ![emyoulation](https://yyz2.discourse-cdn.com/free1/user_avatar/gramps.discourse.group/emyoulation/32/67_2.png) [@emyoulation](https://gramps.discourse.group/u/emyoulation)\
**Post date:** [September 17, 2021, 4:08pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829/3 "2021-09-17T16:08:43Z")

</div>

Thanks for the encouragement. (That’s genuine, not sarcasm.) I’ve been slogging through doing that since posting.

Forked the add-on and doing a first pass as brute force… with a dedicated TSV version that ignores commas as the delimiter. Then I’ll try to figure out how to integrate a delimiter selector.

I _think_ I’ve gotten most of the way to having Gramps recognize a .tsv mime type for import too.

If I get REALLY ambitious, I’ll find a way to have the Text Importer apply a selected Tag or Citation too. I really found those features useful for cleaning up a GEDCOM import. Then discover how to interface that too. (Arrghh!)

---

<div class="post-metadata">

**Author:** ![emyoulation](https://yyz2.discourse-cdn.com/free1/user_avatar/gramps.discourse.group/emyoulation/32/67_2.png) [@emyoulation](https://gramps.discourse.group/u/emyoulation)\
**Post date:** [September 17, 2021, 8:13pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829/4 "2021-09-17T20:13:17Z")

</div>

I did discover that the CSV header labels from the [Export View](https://gramps-project.org/wiki/index.php/Gramps_5.1_Wiki_Manual_-_Settings#Export_View) and what the Import Text will accept are slightly different. (The labels are more compatible when Exporting a Tree than a View.)

Gramps exports views the ID columns labeled as ‘ID’ whereas the [CSV exporter, importer](https://gramps-project.org/wiki/index.php/Gramps_5.1_Wiki_Manual_-_Manage_Family_Trees:_CSV_Import_and_Export) and Import Text Gramplet all expect the Column label for ID to be labeled with the primary object type (Person, Marriage, Event) instead.

I wonder which should be the preferred. But it seems like they should be compatible.

---

<div class="post-metadata">

**Author:** ![emyoulation](https://yyz2.discourse-cdn.com/free1/user_avatar/gramps.discourse.group/emyoulation/32/67_2.png) [@emyoulation](https://gramps.discourse.group/u/emyoulation)\
**Post date:** [January 7, 2023, 7:01pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829/7 "2023-01-07T19:01:43Z")

</div>

Patching eight 5.1.x files with Serge’s changes for Gramps 5.2 gave the TSV functionality for the Text Import Gramplet and CSV Import.

See:

> [@CSV import problem](https://gramps.discourse.group/t/csv-import-problem/2112/4):
>
> Excel was doing bad things to my data. Gramps names & places can include quotes & commas. Either the Excel export was not adding the right Escape codes or the Gramps import wasn’t parsing it correctly. Ether way, CSV was not working well for me.
> 
> @SNoiraud just did [enhancements for Tab Separated Values](https://github.com/gramps-project/gramps/pull/1314) instead using Commas. (Thanks Serge !)
> 
> The enhancements are targeted for release with the 5.2 version. But you may want to try patching a 5.1.x version.

---

<div class="post-metadata">

**Author:** ![emyoulation](https://yyz2.discourse-cdn.com/free1/user_avatar/gramps.discourse.group/emyoulation/32/67_2.png) [@emyoulation](https://gramps.discourse.group/u/emyoulation)\
**Post date:** [September 2, 2024, 3:36pm UTC](https://gramps.discourse.group/t/tab-separated-value-parser-to-supplement-csv/1829/9 "2024-09-02T15:36:24Z")

</div>

The labelings on the CSV dialect options are ambiguous to me.

 
 - excel
 - excel-tab
 - unix
 - Custom  
– ,  
– ;  
– :  
– |  
– Tab  
 

Here are some more explicit labels. But are they correct/accurate? And can we use a MS branded product name without Trademark infringement?

 
 - Excel comma separated values (CSV)
 - Excel tab separated values (TSV)
 - Unix (CSV with LF end-of-line)
 - Custom  
– , (Comma; aka CSV)  
– ; (Semi-colon)  
– : (Colon)  
– | (Vertical Bar aka pipe-delimited)  
– Tab (TSV)  
 

 ![image](https://global.discourse-cdn.com/free1/uploads/gramps/original/2X/5/5c7ab44d1bc633ead3746a53731488f0a0bae003.png)

As for the Unix option, here’s what **Perplexity ai** describes as the difference:

 

> The main differences between a CSV file created by Microsoft Excel on Windows and a Unix CSV file are:
> 
> 1. Line endings:
> 
> - Windows Excel CSV files typically use CRLF (\r\n) for line endings
> - Unix CSV files typically use just LF (\n) for line endings
> 
> 1. Default field delimiter:
> 
> - Windows Excel often uses semicolon ( ; ) as the default field delimiter, especially in regions that use comma as the decimal separator
> - Unix CSV files typically use comma ( , ) as the field delimiter
> 
> 1. Character encoding:
> 
> - Windows Excel CSV files are often saved with Windows-1252 or UTF-8 with BOM encoding
> - Unix CSV files typically use UTF-8 without BOM encoding
> 
> 1. Quoting:  
> Excel may add quotes around fields containing commas or line breaks inconsistently  
> Unix CSV files tend to follow RFC 4180 more strictly for quoting
> 2. Handling of special characters:
> 
> - Excel may handle certain special characters differently than standard Unix tools
> 
> 1. Metadata:
> 
> - Excel CSV files may contain hidden metadata or formatting information
> - Unix CSV files are typically plain text without extra metadata
> 
> 1. Decimal separator:
> 
> - Excel CSV files may use comma or period as decimal separator depending on regional settings
> - Unix CSV files typically use period as decimal separator
> 
> To ensure compatibility when working with CSV files across platforms, it’s often recommended to use a standardized format like RFC 4180 and explicitly specify encoding, delimiters, and line endings when creating or processing CSV files.
