I am trying to extract a news article from a web page. The sed I have written brings back a lot of Javascript code and sometimes advertisments too. Can anyone please help with this one ??? I need to fix this sed so it picks up the article ONLY (don't worry about the title or date .. i got those using a separate sed) ..
hi guys,
am required to prepare a report and mail it, to make it more appealing :p i wish to have content of mail in rich text format i.e html type
with mailx how to specify the content type of mail body as html?
Thanks in advance!!!
rishi (2 Replies)
I have a html file called myfile. If I simply put "cat myfile.html" in UNIX, it shows all the html tags like <a href=r/26><img src="http://www>. But I want to extract only text part.
Same problem happens in "type" command in MS-DOS.
I know you can do it by opening it in Internet Explorer,... (4 Replies)
Hi All,
I'm trying to extract some floating point numbers from within some HTML code like this:
<TR><TD class='awrc'>Parse CPU to Parse Elapsd %:</TD><TD ALIGN='right' class='awrc'> 64.50</TD><TD class='awrc'>% Non-Parse CPU:</TD><TD ALIGN='right' class='awrc'> ... (2 Replies)
Hello,
i try to extract urls from google-search-results, but i have problem with sed filtering of html-code.
what i wont is just list of urls thay apears between ........<p><a href=" and next following " in html code.
here is my code, i use wget and pipelines to filtering. wget works, but... (13 Replies)
I am attempting to extract weather data from the following website, but for the Victoria area only:
Text Forecasts - Environment Canada
I use this:
sed -n "/Greater Victoria./,/Fraser Valley./p"
But that phrasing does not sometimes get it all and think perhaps the website has more... (2 Replies)
Hi
I've searched for it for few hours now and i can't seem to find anything working like i want. I've got webpage, saved in file par with form like this:
<html><body><form name='sendme' action='http://example.com/' method='POST'>
<textarea name='1st'>abc123def678</textarea>
<textarea... (9 Replies)
Hi Expert,
Is there any other way to print and write to a same filename the content between two html tags?
Here the sample:
cat file.html
<div id="outline">
hello world<br>
</div>
<div id="container_faq">
test1<br>
</div>
<div class="widget_quick">
thead test<br>
</div>
... (3 Replies)
I'm extracting text between table tags in HTML
<th><a href="/wiki/Buick_LeSabre" title="Buick LeSabre">Buick LeSabre</a></th>
using this:
awk -F "</*th>" '/<\/*th>/ {print $2}' auto2 > auto3
then this (text between a href):
sed -e 's/\(<*>\)//g' auto3 > auto4
How to shorten this into one... (8 Replies)
Hi, Please see my code below i'm trying get an email send with attachment and html content in the body.
Using the code below will put the encoding for attachment in the body as well
SUBJECT="$(echo "XPI Monitoring "${tcnt}" transactions waiting \nContent-Type: text/html")"
cat... (3 Replies)
Hi
I have file like this:
jack black 104
daniel nick 75
lily harm 2
albert 5
and need to convert it into the html table like this:
NO.......name....family..... id
1...........jack.....black.....104
2..........daniel....nick.......75
3..........albert.................5
i mean... (5 Replies)
Discussion started by: indeed_1
5 Replies
LEARN ABOUT DEBIAN
history
HISTORY(5) File Formats Manual HISTORY(5)NAME
history - record of current and recently expired Usenet articles
DESCRIPTION
The file /var/lib/news/history keeps a record of all articles currently stored in the news system, as well as those that have been received
but since expired. In a typical production environment, this file will be many megabytes.
The file consists of text lines. Each line corresponds to one article. The file is normally kept sorted in the order in which articles
are received, although this is not a requirement. Innd(8) appends a new line each time it files an article, and expire(8) builds a new
version of the file by removing old articles and purging old entries.
Each line consists of two or three fields separated by a tab, shown below as :
<Message-ID> date
<Message-ID> date files
The Message-ID field is the value of the article's Message-ID header, including the angle brackets.
The date field consists of three sub-fields separated by a tilde. All sub-fields are the text representation of the number of seconds
since the epoch -- i.e., a time_t; see gettimeofday(2). The first sub-field is the article's arrival date. If copies of the article are
still present then the second sub-field is either the value of the article's Expires header, or a hyphen if no expiration date was speci-
fied. If an article has been expired then the second sub-field will be a hyphen. The third sub-field is the value of the article's Date
header, recording when the article was posted.
The files field is a set of entries separated by one or more spaces. Each entry consists of the name of the newsgroup, a slash, and the
article number. This field is empty if the article has been expired.
For example, an article cross-posted to comp.sources.unix and comp.sources.d that was posted on February 10, 1991 (and received three min-
utes later), with an expiration date of May 5, 1991, could have a history line (broken into two lines for display) like the following:
<312@litchi.foo.com> 666162000~673329600~666162180
comp.sources.unix/1104 comp.sources.d/7056
In addition to the text file, there is a dbz(3z) database associated with the file that uses the Message-ID field as a key to determine the
offset in the text file where the associated line begins. For historical reasons, the key includes the trailing byte (which is not
stored in the text file).
HISTORY
Written by Rich $alz <rsalz@uunet.uu.net> for InterNetNews. This is revision 1.12, dated 1996/09/06.
SEE ALSO dbz(3z), expire(8), innd(8), news-recovery(8).
HISTORY(5)