Somewhere in Settings for Crawling Fetcher, shouldn't there be a field for start URL? How does the crawler know where to start?

Comments

twistor’s picture

Status: Active » Postponed (maintainer needs more info)

What method are you trying to use?

socratesone’s picture

Issue summary: View changes

I'm running into the same issue, but I'm using 7.x-2.x-dev (along with Feeds 7.x-2.0-alpha8).

I'm using the Feeds interface, selecting "Crawling Fetcher (XPath)" as Fetcher.
In the "settings" page for the Fetcher, there is no field for a url to fetch the data - just the XPath expression. I don't seem to be able to find a url field in the interface, and the documentation available doesn't explain this (I think this is due to the fact that the documentation was written for v 6.x. I thought it might be on the importer form (when you actually import the data), but got the error "Unable to interpret Feeds importer code."

Correct me if I'm wrong, but shouldn't the interface start with the url, then an XPath expression to generate the links, then several other XPath expressions to grab fields from that page? Shouldn't each field in the field mappings be an XPath expression mapped to a field in the content type?

You asked "what method"? Is there any other method here? It would be nice if alternative methods were documented somewhere, too.

ibraaheem’s picture

+1

twistor’s picture

Status: Postponed (maintainer needs more info) » Closed (fixed)

You're either looking for the source url, or the 'URL REPLACEMENT OPTIONS' section.

This version is going to be dead soon. The next one will hopefully make this much simpler.