A Novel Traffic Analysis for Identifying Search Fields in the Long Tail of Web Sites

 
Thread Tools Search this Thread
Special Forums News, Links, Events and Announcements UNIX and Linux RSS News A Novel Traffic Analysis for Identifying Search Fields in the Long Tail of Web Sites
# 1  
Old 02-22-2010
A Novel Traffic Analysis for Identifying Search Fields in the Long Tail of Web Sites

HPL-2010-27 A Novel Traffic Analysis for Identifying Search Fields in the Long Tail of Web Sites - Forman, George; Kirshenbaum, Evan; Rajaram, Shyamsundar
Keyword(s): web data mining, clickstream analysis, machine learning classification, active learning
Abstract: Using a clickstream sample of 2 billion URLs from many thousand volunteer Web users, we wish to analyze typical usage of keyword searches across the Web. In order to do this, we need to be able to determine whether a given URL represents a keyword search and, if so, which field contains the query. A ...
Full Report

More...
Login or Register to Ask a Question

Previous Thread | Next Thread

7 More Discussions You Might Find Interesting

1. What is on Your Mind?

Your Favorite Tech Support Web Sites and Why?

Where do you go to participate in technical discussions besides UNIX.COM and why? Personally, I do not really participate in other forums and discussion boards, but I do ask questions from time to time on Stack sites. The problem I have with Stack is that my questions are never answered on any... (30 Replies)
Discussion started by: Neo
30 Replies

2. Red Hat

Web sites

Hi, I can't view web portal in my intranet from linux RHE, and neither to web application. My network configuration /etc/sysconfig/network-scripts/fcfg-eth0 is ok, what is happen?, can you help me please. (2 Replies)
Discussion started by: xochitl
2 Replies

3. Shell Programming and Scripting

Identifying entries based on 2 fields in a string.

Hi Guys, I’m struggling to use two fields to do a duplicate/ unique by output. I want to look IP addresses assigned to more than one account during a given period in the logs. So duplicate IP and account > 1 then print all the logs for that IP. I have been Using AWK (just as its installed... (3 Replies)
Discussion started by: wabbit02
3 Replies

4. Shell Programming and Scripting

Identifying specific fields in a Row

Hi, I am new to UNIX. Can some one help me to solve the below. I have a requirement to to identify the specific fields in row and also some part of the field. In my file I have a record as sundra;10.44.48.65;10thstreet TCP packet out of state: First packet isn't SYN;telno:... (3 Replies)
Discussion started by: suneel.mekala
3 Replies

5. Web Development

How do you make web sites?

:confused: I've read how on some websites but I still don't get it. I need specific details. I want to make a website for my photography. Please help!:D (3 Replies)
Discussion started by: animelibara123
3 Replies

6. OS X (Apple)

Use UNIX to track web sites viewed?

I'm on OSX 10.4. I was wondering if you can use UNIX terminal to track what web sites have been viewed on this Mac... Thank you! (1 Reply)
Discussion started by: tracymanusa
1 Replies

7. Solaris

Identifying new fields of data

i have hundreds of lines of formatted data with 10 different fields per line. the data is refreshed every few minutes and some fields in some lines may reflect new data. i'm looking for a sample of code that help me to identify those new fields so that i can write them to a file to indicate that... (0 Replies)
Discussion started by: davels
0 Replies
Login or Register to Ask a Question
TM::Analysis(3pm)					User Contributed Perl Documentation					 TM::Analysis(3pm)

NAME
TM::Analysis - Topic Maps, analysis functions SYNOPSIS
use TM::Materialized::AsTMa; my $tm = new TM::Materialized::AsTMa (file => 'test.atm'); $tm->sync_in; Class::Trait->apply ($tm, 'TM::Analysis'); print Dumper $tm->statistics; print Dumper $tm->orphanage; DESCRIPTION
This package contains some topic map analysis functionality. INTERFACE
statistics This (currently quite limited) function computes a reference to hash containing the following fields: "nr_toplets" Nr of midlets in the map. This includes ALL midlets for topics and also those for assertions. "nr_asserts" Nr of assertions in the map. "nr_clusters" Nr of clusters according to the "cluster" function elsewhere in this document. orphanage This computes all topics which have either no supertype and also those which have no type. Without further parameters, it returns a hash reference with the following fields: "untyped" Holds a list reference to all topic ids which have no type. "empty" Holds a list reference to all topic ids which have no instance. "unclassified" Holds a list reference to all topic ids which have no superclass. "unspecified" Holds a list reference to all topic ids which have no subclass. Optionally, a list of the identifiers above can be passed in so that only that particular information is actually returned (some speedup): my $o = TM::Analysis::orphanage ($tm, 'untyped'); entropy This method returns a hash (reference) where the keys are the assertion types and the values are the individual entropies of these assertion types. More frequently used (inflationary) types will have a lower value, very seldomly used ones too. Only those in the middle count most. SEE ALSO
TM COPYRIGHT AND LICENSE
Copyright 20(0[3-68]|10) by Robert Barta, <drrho@cpan.org> This library is free software; you can redistribute it and/or modify it under the same terms as Perl itself. perl v5.10.1 2010-06-06 TM::Analysis(3pm)