Posts

Showing posts from September, 2013

BigData: Apache Flume, HDFS and HBase

Image
In this post, I will show how to log very large amount of web requests to a BigData storage for traffic analysis. The source code for the project is on github . We will rely on the logging library log4j and the associated  Flume NG  appender  implementation . For storage, we will place the log information into a set of  HDFS  files or into an  HBase  table. The HDFS files will be mapped into a Hive table as partitions for query and aggregation. In this demonstration, all web requests are handled by a very simple web application  servlet that is mapped to an access url: In practice, all logic is handled by the servlet and the logging of the request information will be handled by a servlet filter that uses the log4j logging API. In this demonstration the path info will define the log level and the query string will define the log message content. Log4j is configured using the resource log4j.properties to use a Flume NG appender: The ...

Apache HBase Certification

Image
I am now a Cloudera Certified Specialist in Apache HBase . Woohoo !!!

BigData GeoEnrichment

What is GeoEnrichment? An example would best describe it. Given a big set of customer location records, I would like each location to be GeoEnriched with the average income of the zip code where that location falls into and with the number of people between the age of 25 and 30 that live in that zip code. Before GeoEnrichment: CustId,Lat,Lon After GeoEnrichment: CustId,Lat,Lon, AverageIncome,Age25To30 Of course the key to this whole thing is the spatial reference data :-) and there are a lot of search options, such as Point-In-Polygon, Nearest Neighbor and enrichment based on a Drive Time Polygon from each location. I've implemented two search methods: Point-In-Polygon method Nearest Neighbor Weighted method The Point-In-Polygon (PiP) method is fairly simple. Given a point, find the polygon it falls into and pull from the polygon feature the selected attributes and add them to the original point. The Nearest Neighbor Weighted (NNW) method finds all the reference...