organizing

How do you organize your genealogy research?

An illustration of a desk seen from above, with a stack of unsorted papers on the left dissolving into a tidy constellation of linked note cards on the right, lines connecting them.
From a pile to a map. Every record gets one home, and every fact gets a link.

Organize your genealogy research around documents and questions, not around software. Give every record one home and one honest file name. Keep a dated log of what you searched. Tag each fact by how sure you are of it. Let the tree program be the last place a fact goes, not the first. Everything past that is optional detail.

I'll show the system I use on File 001, the Padre Ballí case, because it's public and you can check it. The same four rules work in a shoebox.

Why does organizing feel harder than the research?

Because most people try to organize tools instead of evidence. A thread on r/Genealogy opened with someone who loved the research but spent hours searching for the right way to keep research logs, spreadsheets, note apps and filing systems, then gave up and went to bed. The best replies were not about software. One said to work one focused question at a time, such as whether a particular John is the father of a particular Mary, list the records that could answer it, cite each as you go, and note where you stopped. Others said to try a new system on two or three ancestors first, and to give whatever you pick two weeks before changing it. One more said the only right way is the one that works for you.

I agree with all of them. A record has three things worth keeping together: where it came from, what it says, and what it proves. Any system that holds those three in one place works. Any system that separates them falls apart once the records number in the hundreds. If you're still at the beginning, how to start family history research for free covers the first records to pull. This post is about what to do with them once they're on your drive.

What should my folders and file names look like?

Fewer folders than you think. A case file of mine has six: the case itself (the question, the people, the research paths), the evidence (every scanned record), the letters and corrections sent to archives, the research log and sources, the activity log, and video and publication. The last one exists because File 001 is being filmed.

File names matter more than folders, because a name follows the file wherever it's copied. The r/Genealogy thread on naming systems drew 44 comments and more than a dozen systems. Who-what-when, meaning a surname, a given name, a record type and a date. Date first, so a folder sorts itself into a timeline. Person folders prefixed with a birth date. One person keeps FamilySearch's own digital folder and image numbers as the file name so the page can be found again. Where they overlapped was on three points. Be consistent. Put a date or a number where the computer sorts. File women under the name they were born with.

Mine puts the source first, then the page, then what the page proves. Three real file names from the Ballí evidence folder:

  • DGS004563848_img1240_KEY-balli-burial-16apr1829-finado-11apr.jpeg
  • GLO-SC-000136-2A-p05_will-aug-1828-opening-names-parents.jpg
  • REYNOSA-marriage-reg-seq156_KEY-BALLI-SIGNATURE-jan-feb-1807.jpg

DGS is the number FamilySearch gives a digitized film, and img is the image inside it. GLO SC is a call number, the archive's own shelf code, in the Texas General Land Office's Spanish Collection, and p the page. The third is a sequence number in the Reynosa marriage register as the UTRGV library digitized it. The first half of each name is the citation; the second half is the finding; KEY marks a page that settled something. Sorted by name, the files group by archive, and the first half tells anyone which book to open.

This is a source-first system, suited to a case built from a handful of archives. If you file by person, the birth-date prefix from that thread does the same job. One commenter whose ancestors were Scandinavian warned that surname folders fail where surnames change every generation. Pick the shape of your system from the shape of your records, not from mine.

Do I need a research log, a spreadsheet, or both?

A research log is a dated list of what you searched, where, and what you found, including nothing. The Genea-Musings post of advice for beginning genealogists, from 2015, puts research logs and to-do lists on the list, beside folders by surname or locality for paper, digital folders by family, and a genealogy program for names, dates, places and sources.

The log turns a dead end into a result. On File 001 the public case log is dated and sourced entry by entry, dead ends included. The 132-page island file at the Land Office was read cover to cover to confirm no other death record exists, and that reading is logged with its call number on the case page.

Spreadsheets are a different tool. A thread on r/Genealogy asked what people's spreadsheets look like, and the useful ones were narrow. One was a census grid: a row per ancestor, a column per census year, cells shaded for the years the person was unlikely to appear, and the age and place written into each cell. A wrong match shows up as an age that jumps. Others tracked the document checklist for a citizenship application, a table of DNA matches, and an index to a printed tree. One person trying to sort through a spouse's spreadsheet-as-family-tree found it anything but intuitive. A spreadsheet is for a question that has rows and columns. Everything else is a note.

What do I do with paper, certificates and photos?

In a newcomer's thread on r/Genealogy about physical copies, the two longest replies agreed. Keep paper only for originals and the certified copies you paid an office for. Keep the rest as digital images in folders you control, not only on a subscription site, and back up both locally and to the cloud.

Same here. Paper is for originals, and the scan gets the same file name as every other record. A binder is an output, something to page through, not a filing system. Photos get a date in the name, or a range of years when the exact date is unknown. One commenter in the spreadsheet thread, who names every scan that way, called it honest and accurate. Keep living people's photos out of anything you publish; the tree screenshots on my own site have the living generations cropped out for that reason.

How is the Ballí case file organized?

The vault behind File 001, the case file kept as a set of linked notes, is on the deliverables page as screenshots of the real thing, not a mockup.

One note per document

Every scanned record gets its own note: the citation at the top, the image embedded under it, and a transcription of the lines that matter. The deliverables page shows the note for the January 1807 page where Padre Ballí signed a marriage as interim priest, with the six other places in the file that link to it listed down the side.

Three tags, nothing else

Each fact carries one of three tags: verified, lead, or dead end. Verified means I hold a primary record, one made at the time by someone who was there, that says it. Lead means someone reported it and it's plausible but unproven. Dead end means it was checked and closed. The known-facts note is headed "only what's proven"; leads live in the notes for the research paths that will test them, each keeping its tag. Facts copied from other people's trees come in as leads, never as verified, for the reasons in can you trust other people's family trees.

The tags move. The name of the Padre's grandfather sat in his mother's 1798 will at the Land Office. When the researcher who indexed the Reynosa registers named the same man unprompted on a call, two independent sources agreed, and the fact moved to verified. The 1811 will and the shipwreck came from the same call and went in as leads; the Texas Supreme Court's 1944 opinion in State v. Balli has since put both in the record. The burial entry and the 1828 will were each read for his age; neither states one, so both are logged as dead ends for that question, and the hunt moved to the ordination file in Monterrey.

Backlinks between the will, the survey and the burial

A backlink is a link that runs both ways: open the will and you see every note that cites it. The August 1828 will, the island survey file of 1827 and 1828, and the April 1829 burial entry are linked to one another. That is how the death-date error was caught. The burial register calls him deceased five days before the date every history prints, and the will written eight months earlier has him in full health. Linked, the two could be read against each other instead of one at a time. The graph view on the deliverables page is the same file drawn as a map, each dot a note and each line a citation. The case board is the question in the middle with the research paths branching off it.

The finished form is the family website, where every fact wears a badge: verified, published, reported, or unknown. The sample family site is built from File 001's public records, so you can click through a real one, and the deliverables page has the vault screenshots.

How much of this should you do yourself?

All of it, if you want to. Nothing above needs more than folders, a text editor and a scanner, and a file of a few hundred records with a working tree does not need me. Do the weekly routine in the FAQ below and it will stay organized.

Paying for it makes sense when the file has outgrown the habit: thousands of unnamed downloads, copied trees mixed with proven facts with no way to tell which is which, or a family that wants the file built this way from the first record. The case vault and private website are included from the 1800s package on the pricing page, from $2,500 with $625 to start; on the 1900 package the vault is an add-on from $250 and the website from $300. What I can't do is keep your file alive after delivery. A vault nobody adds to goes stale the way a PDF does, which is why the website exists: one place for the whole family to look and to add to.

If you'd rather have this done, start with the $300 Discovery assessment; it's credited in full toward any package. You get a brick-wall check and a written research plan before you commit to anything larger. Start with the assessment.

Questions people also ask

What is a simple file-naming system for genealogy?

Put the source first, then the page or image number, then a few words on what the page proves, so the citation and the finding travel together. If you file by person instead, lead with the year so each folder sorts itself into a timeline. Whichever you pick, use it on every file, put a date or number where the computer sorts on it, and file women under the name they were born with.

Should I keep a research log?

Yes. A research log is a dated list of what you searched, where, and what you found, including nothing. It is the only thing that stops you, or whoever inherits the file, from repeating a search that already came up empty. One line per session is enough if the line names the collection, the date and the result.

How do I organize paper certificates and photos?

Keep paper for originals and certified copies, in one binder per family with the documents in sleeves. Scan everything, name the scan the way you name every other record, and back the files up in two places. Photos get a date in the file name, or a range of years when the exact date is unknown, plus the names of the people in them, and photos of living relatives stay out of anything you share.

Which genealogy software lets me own my data?

I don't rank products, and the one you will keep using matters more than the brand, so apply one test before you commit. Can you export everything, including sources, notes and images, to files on your own drive, and can you read those files without the program? GEDCOM, the plain-text format tree programs use to exchange data, carries names, dates, relationships and sources, but images travel as separate files, so check what survives a round trip. My own case notes are plain text for the same reason.

What is a 15-minute weekly maintenance routine?

Name and file the records you downloaded this week, so nothing stays in the downloads folder. Add one line to the research log for each search session, including the ones that found nothing. Tag each new fact as verified, lead or dead end, write down the next question, and run your backup. Done every week, that is the whole system.

Sources

  1. Advice for Beginning Genealogists, Genea-Musings (March 2015)
  2. Organizing my research is overwhelming me, a thread on r/Genealogy (December 2021)
  3. Tell me about your genealogy file naming system, a thread on r/Genealogy (December 2022)
  4. For those of you who use spreadsheets to organize your research, a thread on r/Genealogy (April 2023)
  5. Newbie question about organizing info, a thread on r/Genealogy (August 2019)
  6. File 001, The Padre Ballí Birthdate, the public case log
  7. What you actually receive, the deliverables section of the case archive
Keep reading

More from the field notes