Learning scrapers, webcrawlers, search engines and CURL Post: 303019102

Sponsored Content

Operating Systems Linux Learning scrapers, webcrawlers, search engines and CURL Post 303019102 by Neo on Friday 22nd of June 2018 11:16:30 PM

06-23-2018

Administrator

Quote:

Originally Posted by TBotNik

Text only vs regular brower: which is best?
wget vs php fileopen vs CURL: Which is best?
HTML tag find/parse: Are there libraries that effectively do this?
HTML tag find/parse: Is REGEX the best way to parse these? Where are examples?
Checking for the new meta-tags of:

I think you are better off to get web page content using PHP scripts and parse the files with REGEX.

If you Google around, I am sure you can find many sample PHP scripts that do most of what you want. This is very old technology and there is no need to reinvent the wheel parsing HTML data.

Neo

View Public Profile for Neo

Visit Neo's homepage!

Find all posts by Neo

3 More Discussions You Might Find Interesting

1. Shell Programming and Scripting

I dont want to know any search engines

I just want to know where I can download it on this website plz

2. UNIX for Dummies Questions & Answers

Using cURL to save online search results

Hi, I'm attacking this from ignorance because I am not sure how to even ask the question. Here is the mission: I have a list of about 4,000 telephone numbers for past customers. I need to determine how many of these customers are still in business. Obviously, I could call all the numbers....

3. Shell Programming and Scripting

Checking status of engines using C-shell

I am relatively new to scripting. I am trying to develop a script that will 1. Source an executable file as an argument to the script that sets up the environment 2. Run a command "stat" that gives the status of 5 Engines running on the system 3. Check the status of the 5 Engines as either...

LEARN ABOUT PHP

fgetss

FGETSS(3)								 1								 FGETSS(3)

fgetss - Gets line from file pointer and strip HTML tags

SYNOPSIS

       string fgetss (resource	$handle, [int  $length], [string  $allowable_tags])

DESCRIPTION

	Identical to fgets(3), except that fgetss(3) attempts to strip any NUL bytes, HTML and PHP tags from the text it reads.

PARAMETERS

	      o $handle
		-The  file  pointer must be valid, and must point to a file successfully opened by fopen(3) or fsockopen(3) (and not yet closed by
		fclose(3)).

	      o $length
		- Length of the data to be retrieved.

	      o $allowable_tags
		- You can use the optional third parameter to specify tags which should not be stripped.

RETURN VALUES

	Returns a string of up to $length - 1 bytes read from the file pointed to by $handle, with all HTML and PHP code stripped.

	If an error occurs, returns FALSE.

       Example #1

	      Reading a PHP file line-by-line

	      <?php
	      $str = <<<EOD
	      <html><body>
	       <p>Welcome! Today is the <?php echo(date('jS')); ?> of <?= date('F'); ?>.</p>
	      </body></html>
	      Text outside of the HTML block.
	      EOD;
	      file_put_contents('sample.php', $str);

	      $handle = @fopen("sample.php", "r");
	      if ($handle) {
		  while (!feof($handle)) {
		      $buffer = fgetss($handle, 4096);
		      echo $buffer;
		  }
		  fclose($handle);
	      }
	      ?>

	      The above example will output something similar to:

	      Welcome! Today is the  of .

	      Text outside of the HTML block.

NOTES

       Note

	      If PHP is not properly recognizing the line endings when reading files either on or created by a Macintosh  computer,  enabling  the
	      auto_detect_line_endings run-time configuration option may help resolve the problem.

SEE ALSO

       fgets(3), fopen(3), popen(3), fsockopen(3), strip_tags(3).

PHP Documentation Group 														 FGETSS(3)