Usability of modern solutions for logs' analysis and debugging is totally screwed
Why we generate and collect logs? Mostly for further analysis and debugging. For example:
- To find all the error logs with a particular substring in them, and then to inspect them visually.
- To find logs for the particular request_id, user_id or trace_id and then to inspect them visually.
- To calculate the number of successful/unsuccessful hacker attempts to SSH into your host.
- To calculate stats over web logs for a particular ip, domain, etc.
- To calculate the frequency of logs with particular substrings.
All these tasks are easy to perform from command-line when logs are stored in plain files. Just start with cat /path/to/log | grep some-substring. Then iteratively apply the needed commands to the selected logs - wc, awk, more grep, less, head, sort, uniq, cut, etc. - until the desired result is obtained. This approach serves great for analyzing locally stored logs on a few hosts. It doesn't scale well for cases when logs should be analyzed across hundreds of hosts and/or application instances. Of course, there are command-line tools for parallel execution of unix commands across hundreds of hosts, which can help with this case. But we wanted better solution.
So we've got ElasticSearch and Grafana Loki. Both solutions allow collecting logs from hundreds of hosts/applications. But they totally screw up analysis of these logs. They provide awkward to use query languages with silly limitations (such as the number of returned log lines per query) and very limited integration with existing command-line tools for logs' analysis mentioned above. For example, you cannot easily perform the equivalent of cat /log/file | grep some-string | my-custom-script-for-analysis when logs stored in ElasticSearch and Grafana Loki contain millions or billions of lines with some-string substring.
ElasticSearch and Loki also need non-trivial configuration, index creation, performance tuning and maintenance. Do we really want paying this price in exchange to get an awkward ability to analyze logs collected from hundreds of hosts/applications?
Probably, it is time to use better solution, which allows collecting logs from hundreds of sources and then analyzing them with good old command-line tools in the usual ergonomic way? This question was raised many times when I had to analyze logs with modern solutions for logs. I couldn't find the proper solution, so decided creating it on my own based on my experience with creating VictoriaMetrics. So I created open-source user-friendly database for logs - VictoriaLogs. It accepts structured and unstructured logs from popular log shippers such as Filebeat, Fluentbit, Logstash, Vector, etc., it supports fast full-text search without any configuration / tuning, and it has perfect integration with good old command-line tools. Read more about the integration here. Give it a try and share your experience!