Ohi, author here! Thanks for posting Hister. Feel free to A.M.A.
My first free software search project was Searx, a privacy respecting metasearch engine, but because of the limitations of the metasearch concept, I've decided to take a different approach.
Hister builds a personal search index from pages you visit, bookmarks, browser history, local files, and crawled websites. It stores extracted content with offline result previews, so information remains searchable even when the original page changes or disappears. It supports full text and semantic search, can run entirely on your own machine, and includes a web interface, command line tools, and an MCP endpoint for assistant integrations.
Ps.: It looks like our name conflicts with a registered trademark in the US. The owner of the other project has asked us to change it, so we’ll probably need to comply sooner or later.
Name suggestions are welcome! Ideally, the new name should be relatively short, sound good, and have an available .org domain.
The thing I most struggle with in this domain is recalling information from videos. I watch/listen to a lot of hour+ lectures and I rely on this website deeply:
It lets you search YouTube transcripts. If you could somehow integrate video transcripts into this tool, I would be extremely interested in trying it out
Been using hister for a number of weeks now; i'm coming across sites whose content would be better handled with a custom extractor; but it looks like extractors need to be bundled into the build in order to work? Is that correct?
Put another way, i can't write an extractor for Reuters and then point a config to it from my current hister binary?
Hi asciimoo, seems like this is the second time Hister is hitting the HN front page in a month, so congrats on the success!
Question for you: For the less tech savvy of us on here, is there any chance Hister can be can hosted on something like Pikapods? https://www.pikapods.com/
Yes, that's something I'd like to support. The main missing piece for a user-friendly hosting option such as PikaPods is a configuration UI. At the moment, customizing Hister requires editing a configuration file, which isn't practical for this kind of hosted service.
Appreciate the response! I'll be eagerly following Hister's progress. For now, I've settled on a mix of Instapaper and using SingleFile uploads to Dropbox.
Exactly, this is the biggest advantage of the extension. It is fully invisible for the websites, so no captcha, anti-bot protection, no authentication issues, every common bottleneck of a classic crawler is solved by the browser/user.
This is a great suggestion, thanks! I'll definitely add it to the list of candidates. My plan is to do a vote on our social platforms if we have a few decent candidates.
Genuinely curious, how could one fight pre-emptive domain squatters once any candidate is publicly suggested?
When I have suggested names in other situations like this in the past, I spent the ~$10 to get the domain, and offered the transfer the free. Of course, not everyone would do this.
I ran it for most of this year but encountered some problems with it I couldn't fix and thus have not had it hooked up to anything since June when I finally couldn't take it anymore.
In short, I serve a good number of apps from an Nginx reverse proxy. Maybe 25% of them are exposed to the WWW while everything else is limited to the LAN but I still get valid TLS for all of it.
Hister, though, kept breaking my whole reverse proxy and I could never figure out EXACTLY why so I could fix it. After running fine for a few days, it would hog the whole server and everything else proxied by Nginx would become unreachable. I tried tuning the config for it to no avail.
One day when I'm less lazy, I'll probably hook it back up via it's LAN IP to every machine I've got again. I REALLY liked that I could log my browsing history from any machine anywhere in the world without a VPN and I was really disappointed when I had to disable its config in Nginx.
I still use it a lot to go find stuff I flagged as important quickly.
I'm curious if this is something you've heard of before, or if I've got a one off problem here.
I even ported the config to a brand new VM with NGINX and still had the same problem.
Are you sure it was the proxy and not JavaScript in your browser? I've seen similar behaviour from one particular website where using it in a certain way causes it to process a lot of data slowly and block the main thread. This somehow persisted across tabs, even if I closed all the tabs and tried again that site was still hanging until that process finished. But if I used incognito or another browser it would he responsive.
It might be something similar if all of your sites are subdomains. Try incognito at the same time next time
That's a good call out but it is something I tested originally and ruled out.
I've got some applications used daily by friends all over the world and the services would all become unavailable to them when this started happening. Only fix I found was restarting NGINX service and then it could happen again an hour later or 3 days later. Once I removed the proxy config for Hister from the service, the issue never happened again.
I'll repro the issue and get the details intoa. Github Issue this weekend.
Yes, I should have some time this weekend to reproduce the problem and provide useful information for ya'll to look at...even if it's just to determine the problem is for me to fix and not a bug with Hister.
What does Hister do differently? Search seems like a major differentiator, I'm wondering if leveraging the existing archivebox project for archival and implementing good search on top would be more efficient
The main difference I see is Hister focuses on creating an active knowledge base and finding information quickly, while ArchiveBox focuses on preserving web content for the long term.
Thanks for answering. Do you think these dovetail? Both archive everything you browse, so that's common functionality that could be factored out. I only want one archive, having two separate archives because one focuses on search and the other on long term archival is inefficient. What do you do for long term archival - or do you not have this use case?
Thanks so much for creating this. Installed last time it was posted and have been loving it. The MCP server and extensions and userscripts are great QOL additions, as well. Always wondered if something was out there like this and you answered my prayers! New name suggestion: MisterHistory
Would like:
* Local web page interface or even browser UI element (since extension needed anyway)
* Ability to add notes to history
* Flag if bookmarked, allow filtering "bookmarks only"
* Keep old versions of pages
* Human-readable text diff vs current live page
Whatever you do, try not to chose a name which collides with nostrodamus' predictions about .. (ok, he actually used "Hister" to refer to the Danube it seems, but popular legend has another take which is .. unfortunate)
Not a Lawyer, but Hister is a common, old word. Unless you're in the same business as the trademark holder, I would think you have grounds to continue? I note you have the domain name registered, and any deep-pocketed hostile trademark owner would already have reclaimed that. Unfortunately, legal advice is expensive, and the system is open to abuse.
Hey thanks! Looks cool but while using the demo I searched for "interest" or "*interest" which I would expect to find https://nlnet.nl/themes/ which was one of 28 pages with "Areas of special interest" as the H1 title innerText and yet the side panel populated the page for "Internet" (https://en.wikipedia.org/wiki/Internet). I assume there's some fuzzy search stuff going on here any thoughts?
I built something similar at the start of the year, using tailscale for auth (multi user on my tailnet/home network) and for access wherever I was, whether at home or on the road (all compute/storage was on my Mac mini.) Worked a treat.
My suggestion is Lookdown. Kind of a play on words reminiscent of Markup/Markdown --> Lookup/Lookdown. But also a lookdown is a distinctive looking fish with a cute concave profile so you've got an instant mascot.
I gather rss feeds from the websites I visit and it's hard to express how interesting they are. The gut says it borderlines some random collection but that couldn't be more wrong. I also enjoyed YaCy, that project should have a good amount of ideas for you. I kinda end up assigning more and more bandwidth until it gets in the way and I forget to enable it again. The turtle button on some torrent clients is a good invention.
Hyster (not Hister as far as I can tell) is a registered trademark in the US. Regardless:
"The HYSTER trademark is filed in the category of Education and Entertainment Services" [0], so you can safely ignore any demands to rename, as it doesn't conflict. Embarrassing for them that their lawyers don't understand even the basics of trademark law.
Edit to clarify as some folks here are as confused as sHyster's lawyers: Trademarks are not exclusive, they're restricted to a category or categories. You can be Apple in the category of computers, but not in music if there's already an Apple in that category (unless you have $500M to spare [2]).
[0] https://www.trademarkia.com/hyster-77843354
[1] https://tmsearch.uspto.gov/ for a more authoritative source than [0], but it doesn't allow deep linking
[2] https://en.wikipedia.org/wiki/Apple_Corps_v_Apple_Computer
It feels like I'm the only person using this but I'd like to throw another potential bookmark manager integration into the ring, cherry https://github.com/haishanh/cherry
I doubt it. Trademark “infringement “ only happens when the two parties compete in the same space. I think if the other party is the music party game, you’re pretty much in the clear. They may still sue you and lose unless you cave.
Hister builds a personal search index from pages you visit, bookmarks, browser history, local files, and crawled websites. It stores extracted content with offline result previews, so information remains searchable even when the original page changes or disappears. It supports full text and semantic search, can run entirely on your own machine, and includes a web interface, command line tools, and an MCP endpoint for assistant integrations.
Website: https://hister.org/
Tiny read-only demo: https://demo.hister.org/
Ps.: It looks like our name conflicts with a registered trademark in the US. The owner of the other project has asked us to change it, so we’ll probably need to comply sooner or later.
Name suggestions are welcome! Ideally, the new name should be relatively short, sound good, and have an available .org domain.
Thanks!