Thursday, July 23, 2009

Nutch: Custom Plugin to parse and add a field

Last week, I described my initial explorations with Nutch, and the code for a really simple plugin. This week, I describe a pair of plugin components that parse out the blog tags (the "Labels:" towards the bottom of this page) and add them to the index. The plugins are fairly useless in a general case, because you cannot depend on a particular format of a page unless, like me, you are looking at a very small subset of the web. I wrote them in order to understand Nutch's plugin architecture, and to see what was involved in using it as a crawler-indexer combo.

Most of the code in here is based on the information I found in the Nutch Writing Plugin Example wiki page, which is based on Nutch 0.9. There are some API changes between Nutch 0.9 and Nutch 1.0 (which I use), which I had to look at the contributed plugin source code to figure out, but other than that, my example is quite vanilla.

I am adding to the same myplugins plugin that I described in my previous post. My new plugin pair consists of a Parsing filter to parse out the tags from the HTML page, and an Indexing filter to put the tags into the Lucene index. My new plugin.xml file looks like this:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
<?xml version="1.0" encoding="UTF-8"?>
<plugin id="myplugins" name="My test plugins for Nutch"
     version="0.0.1" provider-name="mycompany.com">

   <runtime>
      <library name="myplugins.jar">
         <export name="*"/>
      </library>
   </runtime>

   <extension id="com.mycompany.nutch.indexing.InvalidUrlIndexFilter"
       name="Invalid URL Index Filter"
       point="org.apache.nutch.indexer.IndexingFilter">
     <implementation id="MyPluginsInvalidUrlFilter"
         class="com.mycompany.nutch.indexing.InvalidUrlIndexFilter"/>
   </extension>
   
   <extension id="com.mycompany.nutch.parsing.TagExtractorParseFilter"
       name="Tag Extractor Parse Filter"
       point="org.apache.nutch.parse.HtmlParseFilter">
     <implementation id="MyPluginsTagExtractorParseFilter"
         class="com.mycompany.nutch.parsing.TagExtractorParseFilter"/>
   </extension>
   
   <extension id="com.mycompany.nutch.parsing.TagExtractorIndexFilter"
       name="Tag Extractor Index Filter"
       point="org.apache.nutch.indexer.IndexingFilter">
     <implementation id="MyPluginsTagExtractorIndexFilter"
         class="com.mycompany.nutch.indexing.TagExtractorIndexFilter"/>
   </extension>
</plugin>

The code for the TagExtractorParseFilter is shown below. It reads the content byte array line by line and applies regular expressions to a particular portion of the page to extract tags, then stuff them into a named slot in the parse MetaData map for retrieval and use by the corresponding indexing filter. The class implements the HtmlParseFilter interface, and will be called as one of the configured HtmlParseFilters when the Nutch parse subcommand is run (described below).

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
// Source: src/plugin/myplugins/src/java/com/mycompany/nutch/parsing/TagExtractorParseFilter.java
package com.mycompany.nutch.parsing;

import java.io.BufferedReader;
import java.io.ByteArrayInputStream;
import java.io.IOException;
import java.io.InputStreamReader;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

import org.apache.commons.lang.StringUtils;
import org.apache.hadoop.conf.Configuration;
import org.apache.log4j.Logger;
import org.apache.nutch.metadata.Metadata;
import org.apache.nutch.parse.HTMLMetaTags;
import org.apache.nutch.parse.HtmlParseFilter;
import org.apache.nutch.parse.Parse;
import org.apache.nutch.parse.ParseResult;
import org.apache.nutch.parse.ParseText;
import org.apache.nutch.protocol.Content;
import org.w3c.dom.DocumentFragment;

/**
 * The parse portion of the Tag Extractor module. Parses out blog tags 
 * from the body of the document and sets it into the ParseResult object.
 */
public class TagExtractorParseFilter implements HtmlParseFilter {

  public static final String TAG_KEY = "labels";
  
  private static final Logger LOG = 
    Logger.getLogger(TagExtractorParseFilter.class);
  
  private static final Pattern tagPattern = 
    Pattern.compile(">(\\w+)<");
  
  private Configuration conf;

  /**
   * We use regular expressions to parse out the Labels section from
   * the section snippet shown below:
   * <pre>
   * Labels:
   * <a href='http://sujitpal.blogspot.com/search/label/ror' rel='tag'>ror</a>,
   * ...
   * </span>
   * </pre>
   * Accumulate the tag values into a List, then stuff the list into the
   * parseResult with a well-known key (exposed as a public static variable
   * here, so the indexing filter can pick it up from here).
   */
  public ParseResult filter(Content content, ParseResult parseResult,
      HTMLMetaTags metaTags, DocumentFragment doc) {
    LOG.debug("Parsing URL: " + content.getUrl());
    BufferedReader reader = new BufferedReader(
      new InputStreamReader(new ByteArrayInputStream(
      content.getContent())));
    String line;
    boolean inTagSection = false;
    List<String> tags = new ArrayList<String>();
    try {
      while ((line = reader.readLine()) != null) {
        if (line == null) {
          continue;
        }
        if (line.contains("Labels:")) {
          inTagSection = true;
          continue;
        }
        if (inTagSection && line.contains("</span>")) {
          inTagSection = false;
          break;
        }
        if (inTagSection) {
          Matcher m = tagPattern.matcher(line);
          if (m.find()) {
            LOG.debug("Adding tag=" + m.group(1));
            tags.add(m.group(1));
          }
        }
      }
      reader.close();
    } catch (IOException e) {
      LOG.warn("IOException encountered parsing file:", e);
    }
    Parse parse = parseResult.get(content.getUrl());
    Metadata metadata = parse.getData().getParseMeta();
    for (String tag : tags) {
      metadata.add(TAG_KEY, tag);
    }
    return parseResult;
  }

  public Configuration getConf() {
    return conf;
  }

  public void setConf(Configuration conf) {
    this.conf = conf;
  }
}

The TagExtractorIndexFilter is the other part of this pair. This retrieves the value of the labels from the Parse object and sticks it into the Lucene index. The code is shown below.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
// Source: src/plugin/myplugins/src/java/com/mycompany/nutch/indexing/TagExtractorIndexFilter.java
package com.mycompany.nutch.indexing;

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.io.Text;
import org.apache.log4j.Logger;
import org.apache.nutch.crawl.CrawlDatum;
import org.apache.nutch.crawl.Inlinks;
import org.apache.nutch.indexer.IndexingException;
import org.apache.nutch.indexer.IndexingFilter;
import org.apache.nutch.indexer.NutchDocument;
import org.apache.nutch.indexer.lucene.LuceneWriter;
import org.apache.nutch.indexer.lucene.LuceneWriter.INDEX;
import org.apache.nutch.indexer.lucene.LuceneWriter.STORE;
import org.apache.nutch.parse.Parse;

import com.mycompany.nutch.parsing.TagExtractorParseFilter;

/**
 * The indexing portion of the TagExtractor module. Retrieves the
 * tag information stuffed into the ParseResult object by the parse
 * portion of this module.
 */
public class TagExtractorIndexFilter implements IndexingFilter {

  private static final Logger LOGGER = 
    Logger.getLogger(TagExtractorIndexFilter.class);
  
  private Configuration conf;
  
  public void addIndexBackendOptions(Configuration conf) {
    LuceneWriter.addFieldOptions(
      TagExtractorParseFilter.TAG_KEY, STORE.YES, INDEX.UNTOKENIZED, conf);
  }

  public NutchDocument filter(NutchDocument doc, Parse parse, Text url,
      CrawlDatum datum, Inlinks inlinks) throws IndexingException {
    String[] tags = 
      parse.getData().getParseMeta().getValues(
      TagExtractorParseFilter.TAG_KEY);
    if (tags == null || tags.length == 0) {
      return doc;
    }
    // add to the nutch document, the properties of the field are set in
    // the addIndexBackendOptions method.
    for (String tag : tags) {
      LOGGER.debug("Adding tag: [" + tag + "] for URL: " + url.toString());
      doc.add(TagExtractorParseFilter.TAG_KEY, tag);
    }
    return doc;
  }

  public Configuration getConf() {
    return this.conf;
  }

  public void setConf(Configuration conf) {
    this.conf = conf;
  }
}

We already have the myplugins plugin registered with Nutch, so to exercise the parser, we generate the set of urls to fetch from crawldb, fetch the pages without parsing, then parse. Once that is done, we run updatedb to update the crawldb, then index, dedup and merge. The entire sequence of commands is listed below.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
sujit@sirocco:/opt/nutch-1.0$ CRAWL_DIR=/home/sujit/tmp
sujit@sirocco:/opt/nutch-1.0$ bin/nutch generate \
  $CRAWL_DIR/data/crawldb $CRAWL_DIR/data/segments

# this will create a segments subdirectory which is used in the
# following commands (we set it to SEGMENTS_DIR below)
sujit@sirocco:/opt/nutch-1.0$ SEGMENTS_DIR=20090720105503
sujit@sirocco:/opt/nutch-1.0$ bin/nutch fetch \
  $CRAWL_DIR/data/segments/$SEGMENTS_DIR -noParsing

# The parse command is where our custom parsing happens.
# To run this (for testing) multiple times, remove the 
# crawl_parse, parse_date and parse_text under the
# segments subdirectory after a failed run.
sujit@sirocco:/opt/nutch-1.0$ bin/nutch parse \
  $CRAWL_DIR/data/segments/$SEGMENTS_DIR
sujit@sirocco:/opt/nutch-1.0$ bin/nutch updatedb \
  $CRAWL_DIR/data/crawldb -dir $CRAWL_DIR/data/segments/*

# The index command is where our custom indexing happens
# you should remove $CRAWL_DIR/data/index and 
# $CRAWL_DIR/data/indexes from previous run before running
# these commands.
sujit@sirocco:/opt/nutch-1.0$ bin/nutch index \
  $CRAWL_DIR/data/indexes $CRAWL_DIR/data/crawldb \
  $CRAWL_DIR/data/linkdb $CRAWL_DIR/data/segments/*
sujit@sirocco:/opt/nutch-1.0$ bin/nutch dedup \
  $CRAWL_DIR/data/indexes
sujit@sirocco:/opt/nutch-1.0$ bin/nutch merge \
  -workingdir $CRAWL_DIR/data/work $CRAWL_DIR/data/index \
  $CRAWL_DIR/data/indexes

I tried setting the fetcher.parse value to false in my conf/nutch-site.xml but it did not seem to have any effect - I had to set the -noParsing flag in my nutch fetch command. But here is the snippet, just in case.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
<?xml version="1.0"?>
<?xml-stylesheet type="text/xsl" href="configuration.xsl"?>

<configuration>
  ...
  <property>
    <name>fetcher.parse</name>
    <value>false</value>
  </property>
  ...
</configuration>

I noticed this error message when I was running the parse.

1
2
3
4
5
6
7
Error parsing: http://sujitpal.blogspot.com/feeds/6781582861651651982
/comments/default: org.apache.nutch.parse.ParseException: parser not found
for contentType=application/atom+xml url=http://sujitpal.blogspot.com
/feeds/6781582861651651982/comments/default
 at org.apache.nutch.parse.ParseUtil.parse(ParseUtil.java:74)
 at org.apache.nutch.fetcher.Fetcher$FetcherThread.output(Fetcher.java:766)
 at org.apache.nutch.fetcher.Fetcher$FetcherThread.run(Fetcher.java:552)

I noticed that application/atom+xml was not explicitly mapped to a plugin, so I copied the application/rss+xml setting to it in conf/parse-plugins.xml, and adding parse-rss to plugin.includes in conf/nutch-site.xml. The error message disappeared, but was replaced with a warning about the parse-rss not being able to parse the atom+xml content properly. I did not investigate further, because I was throwing away these pages anyway (InvalidUrlIndexFilter). Here is the snippet from parse-plugins.xml.

1
2
3
4
        <mimeType name="application/atom+xml">
            <plugin id="parse-rss" />
            <plugin id="feed" />
        </mimeType>

You can test (apart from verifying the log traces in hadoop.log) that the TagExtractor combo worked by looking at the Lucene index generated. Here is a screenshot of the top terms for the "label" field.

I probably should have gone all the way and built a QueryFilter to look for this field in the search queries, but I really have no intention of ever using Nutch's search service (except perhaps as an online debugging tool, similar to Luke, which is a lot of work at this point) so I decided not to.

You may have noticed that the code above has nothing to do with Map-Reduce, these are simply (well, almost) plain Java hooks that plugin in to Nutch's published extension points. However, the Nutch core, which calls these plugins during its life-cycle, uses Hadoop Map-Reduce pretty heavily, as these Map-Reduce in Nutch slides by Doug Cutting shows. I plan on looking at this stuff more over the coming week, and possibly write about it if I come up with anything interesting.