Learning scrapers, webcrawlers, search engines and CURL Post: 303019102

Sponsored Content

Operating Systems Linux Learning scrapers, webcrawlers, search engines and CURL Post 303019102 by Neo on Friday 22nd of June 2018 11:16:30 PM

06-23-2018

Administrator

Quote:

Originally Posted by TBotNik

Text only vs regular brower: which is best?
wget vs php fileopen vs CURL: Which is best?
HTML tag find/parse: Are there libraries that effectively do this?
HTML tag find/parse: Is REGEX the best way to parse these? Where are examples?
Checking for the new meta-tags of:

I think you are better off to get web page content using PHP scripts and parse the files with REGEX.

If you Google around, I am sure you can find many sample PHP scripts that do most of what you want. This is very old technology and there is no need to reinvent the wheel parsing HTML data.

Neo

View Public Profile for Neo

Visit Neo's homepage!

Find all posts by Neo

3 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

I dont want to know any search engines

I just want to know where I can download it on this website plz

2. UNIX for Dummies Questions & Answers

Using cURL to save online search results

Hi, I'm attacking this from ignorance because I am not sure how to even ask the question. Here is the mission: I have a list of about 4,000 telephone numbers for past customers. I need to determine how many of these customers are still in business. Obviously, I could call all the numbers....

3. Shell Programming and Scripting

Checking status of engines using C-shell

I am relatively new to scripting. I am trying to develop a script that will 1. Source an executable file as an argument to the script that sets up the environment 2. Run a command "stat" that gives the status of 5 Engines running on the system 3. Check the status of the 5 Engines as either...

LEARN ABOUT PHP

get_meta_tags

GET_META_TAGS(3)							 1							  GET_META_TAGS(3)

get_meta_tags - Extracts all meta tag content attributes from a file and returns an array

SYNOPSIS

       array get_meta_tags (string  $filename, [bool  $use_include_path = false])

DESCRIPTION

	Opens $filename and parses it line by line for <meta> tags in the file. The parsing stops at </head>.

PARAMETERS

	      o $filename
		- The path to the HTML file, as a string. This can be a local file or an URL.

	      Example #1

		     What get_meta_tags(3) parses

		     <meta name="author" content="name">
		     <meta name="keywords" content="php documentation">
		     <meta name="DESCRIPTION" content="a php manual">
		     <meta name="geo.position" content="49.33;-86.59">
		     </head> <!-- parsing stops here -->
	       (pay attention to line endings - PHP uses a native function to parse the input, so a Mac file won't work on Unix).

	      o $use_include_path
		-  Setting  $use_include_path  to  TRUE  will  result  in  PHP	trying to open the file along the standard include path as per the
		include_path directive. This is used for local files, not URLs.

RETURN VALUES

	Returns an array with all the parsed meta tags.

	The value of the name property becomes the key, the value of the content property becomes the value of the returned array, so you can eas-
       ily  use  standard array functions to traverse it or access single values. Special characters in the value of the name property are substi-
       tuted with '_', the rest is converted to lower case. If two meta tags have the same name, only the last one is returned.

EXAMPLES

       Example #2

	      What get_meta_tags(3) returns

	      <?php
	      // Assuming the above tags are at www.example.com
	      $tags = get_meta_tags('http://www.example.com/');

	      // Notice how the keys are all lowercase now, and
	      // how . was replaced by _ in the key.
	      echo $tags['author'];	  // name
	      echo $tags['keywords'];	  // php documentation
	      echo $tags['description'];  // a php manual
	      echo $tags['geo_position']; // 49.33;-86.59
	      ?>

NOTES

       Note

	       Only meta tags with name attributes will be parsed. Quotes are not required.

SEE ALSO

       htmlentities(3), urlencode(3).

PHP Documentation Group 													  GET_META_TAGS(3)