BigData: Apache Flume, HDFS and HBase
In this post, I will show how to log very large amount of web requests to a BigData storage for traffic analysis. The source code for the project is on github . We will rely on the logging library log4j and the associated Flume NG appender implementation . For storage, we will place the log information into a set of HDFS files or into an HBase table. The HDFS files will be mapped into a Hive table as partitions for query and aggregation. In this demonstration, all web requests are handled by a very simple web application servlet that is mapped to an access url: In practice, all logic is handled by the servlet and the logging of the request information will be handled by a servlet filter that uses the log4j logging API. In this demonstration the path info will define the log level and the query string will define the log message content. Log4j is configured using the resource log4j.properties to use a Flume NG appender: The ...