brev
Feed aggregator for RSS, Atom, and pretty much anything else.
I wanted the convenience of following things through RSS, even for things that don’t have RSS feeds. After some previous attempts and further thinking about it, I came up with the idea of something that just fetches the content of a URL, and then passes that to an external parser. The choice of using Python was due to it being easy to just run, and in order to make it as easy as possible I also decided to not use any external dependencies beyond basic Python.
When running the script, it first finds all feed files and loads them in one by one.
def get_feed_paths(top_directory, file_ending):
paths = []
for directory, _, files in os.walk(top_directory):
for file in files:
if file.endswith(file_ending):
paths.append(os.path.join(directory, file))
return paths
def load_feed(path):
with open(path, "r") as file:
try:
lines = file.read().splitlines()
feed = {}
feed["url"] = lines[0]
feed["parser"] = lines[1]
feed["entries"] = lines[2:]
return True, feed
except Exception as e:
return False, e
As you can see, the feed files have a very simple format: the first line is the URL, the second line is the parser, and the remaining lines are entries in the feed. It is up the parsers what entries look like, as long as they’re one line each.
Another reason why I used Python is that the parsers can be external scripts and get imported dynamically without needing to be registered anywhere.
def import_parsers():
parsers = {}
for file in os.listdir():
if file.endswith(".py") and file != sys.argv[0]:
name = file[:-len(".py")]
parsers[name] = importlib.import_module(name)
return parsers
This makes it easy to add a simple parser for one specific thing.
# Parser for specifically https://data.iana.org/TLD/tlds-alpha-by-domain.txt
# Takes a string and returns a list of strings
def parse(raw_feed):
return raw_feed.splitlines()[1:] # Every line except the first
For each feed, the content at the URL is fetched:
def fetch_raw_feed(url):
try:
request = urllib.request.Request(url)
request.add_header("User-Agent", "brev")
with urllib.request.urlopen(request) as response:
if response.status == 200:
encoding = response.headers.get_content_charset()
if encoding is None:
encoding = "utf-8"
return True, response.read().decode(encoding)
else:
return False, response.status
except Exception as e:
return False, e
And then passed to a parser:
if feed["parser"] in parsers:
try:
entries = parsers[feed["parser"]].parse(raw_feed)
except Exception as e:
out += error_log(f"FAILED TO PARSE FEED '{feed_name}' WITH PARSER '{feed['parser']}' ({e})")
continue
else:
entries = regex_parse(raw_feed, feed["parser"])
If you don’t want to create a specific parser for something, you can specify a regex string on the second line instead. The built-in regex parser then returns every line matching that regex as an entry in the feed.
def regex_parse(raw_feed, pattern_string):
entries = []
pattern = re.compile(pattern_string)
for line in raw_feed.splitlines():
if pattern.search(line):
entries.append(line.strip())
return entries
Finally, the new entries get added to the output file and the current state of the feed gets saved.
def get_new_entries(entries, old_entries, feed_name, output_format):
out = ""
for entry in entries:
if entry not in old_entries:
out += output_format.format(feed_name = feed_name, entry = entry) + "\n"
return out
def save_feed(path, url, parser, entries):
with open(path, "w") as file:
file.write("{url}\n{parser}\n{entries}\n".format(url = url, parser = parser, entries = "\n".join(entries)))
With this way of doing things even parsing RSS feeds isn’t a built-in feature, but that just means everyone can do it how they prefer. Personally I only care about titles and links.
import xml.etree.ElementTree as et
# Parser for RSS and Atom
# Takes a string and returns a list of strings
def parse(raw_feed):
root = et.fromstring(raw_feed)
ns = {}
if root.tag == "{http://www.w3.org/1999/02/22-rdf-syntax-ns#}RDF": # RSS 1.0
ns[""] = "http://purl.org/rss/1.0/"
items = root.findall("item", namespaces = ns)
elif root.tag == "rss": # RSS 2.0
items = root.find("channel").findall("item")
elif root.tag == "{http://www.w3.org/2005/Atom}feed": # Atom
ns[""] = "http://www.w3.org/2005/Atom"
items = root.findall("entry", namespaces = ns)
else:
raise RuntimeError(f"root tag of unknown type '{root.tag}'")
entries = []
for item in items:
title = item.findtext("title", namespaces = ns) # RSS and Atom
link = item.findtext("link", namespaces = ns) # RSS
if link == "" or link is None:
linktag = item.find("link", namespaces = ns)
enclosuretag = item.find("enclosure", namespaces = ns)
if linktag is not None:
link = linktag.attrib["href"] # Atom
elif enclosuretag is not None:
link = enclosuretag.attrib["url"] # Podcasts without link tags, get link to audio/video file instead
entries.append(" ".join(f"{title} {link}".splitlines()))
return entries