I am building an Online news clustering system using Lucene and Mahout libraries in java. I intend to use vector space model and tfidf weights for Kmeans(or fuzzy/streamKmeans). My plan is : Cluster initial articles,assign new article to the cluster whose centroid is closest based on a small distance threshold. The leftover documents that aren’t associated with any old clusters form new data(new topics). Separately cluster them among themselves and add these temporary cluster centroids to the previous centroids. Less frequently, execute the full batch clustering to recluster the entire set of documents. The problem arises in comparing a new article to a centroid to assign it to an old cluster. The centroid dimension is number of distinct words in initial data. But the dimension of new article is different. I am following the book Mahout in Action. Is there any approach or some sort of feature extraction to handle this. The following similar links still remain unanswered: https://stats.stackexchange.com/questions/41409/bag-of-words-in-an-online-configuration-for-classification-clustering https://stats.stackexchange.com/questions/123830/vector-space-model-for-online-news-clustering Thanks in advance
Incorporating new articles in tfidf vector for online clustering
104 views Asked by aman2357 At
1
There are 1 answers
Related Questions in CLUSTER-ANALYSIS
- google speech recognition api in hindi
- Unable to copy exact hindi content from pdf
- Change locale in android app (onto Hindi)
- Weird characters while displaying hindi text in android
- BreakIterator in Android counts character wrongly
- How to detect if a string contains hindi (devnagri) in it with character and word count
- Tcpdf Hindi sentence display issue
- How to set UTF-8 encoding in solr-php client to store hindi content in solr
- I want to upload file containing Hindi language words and English words to phpMyadmin but getting errors
- AS3 - Hindi font half letters getting converted
Related Questions in MAHOUT
- google speech recognition api in hindi
- Unable to copy exact hindi content from pdf
- Change locale in android app (onto Hindi)
- Weird characters while displaying hindi text in android
- BreakIterator in Android counts character wrongly
- How to detect if a string contains hindi (devnagri) in it with character and word count
- Tcpdf Hindi sentence display issue
- How to set UTF-8 encoding in solr-php client to store hindi content in solr
- I want to upload file containing Hindi language words and English words to phpMyadmin but getting errors
- AS3 - Hindi font half letters getting converted
Related Questions in K-MEANS
- google speech recognition api in hindi
- Unable to copy exact hindi content from pdf
- Change locale in android app (onto Hindi)
- Weird characters while displaying hindi text in android
- BreakIterator in Android counts character wrongly
- How to detect if a string contains hindi (devnagri) in it with character and word count
- Tcpdf Hindi sentence display issue
- How to set UTF-8 encoding in solr-php client to store hindi content in solr
- I want to upload file containing Hindi language words and English words to phpMyadmin but getting errors
- AS3 - Hindi font half letters getting converted
Related Questions in TEXT-MINING
- google speech recognition api in hindi
- Unable to copy exact hindi content from pdf
- Change locale in android app (onto Hindi)
- Weird characters while displaying hindi text in android
- BreakIterator in Android counts character wrongly
- How to detect if a string contains hindi (devnagri) in it with character and word count
- Tcpdf Hindi sentence display issue
- How to set UTF-8 encoding in solr-php client to store hindi content in solr
- I want to upload file containing Hindi language words and English words to phpMyadmin but getting errors
- AS3 - Hindi font half letters getting converted
Related Questions in TF-IDF
- google speech recognition api in hindi
- Unable to copy exact hindi content from pdf
- Change locale in android app (onto Hindi)
- Weird characters while displaying hindi text in android
- BreakIterator in Android counts character wrongly
- How to detect if a string contains hindi (devnagri) in it with character and word count
- Tcpdf Hindi sentence display issue
- How to set UTF-8 encoding in solr-php client to store hindi content in solr
- I want to upload file containing Hindi language words and English words to phpMyadmin but getting errors
- AS3 - Hindi font half letters getting converted
Popular Questions
- How do I undo the most recent local commits in Git?
- How can I remove a specific item from an array in JavaScript?
- How do I delete a Git branch locally and remotely?
- Find all files containing a specific text (string) on Linux?
- How do I revert a Git repository to a previous commit?
- How do I create an HTML button that acts like a link?
- How do I check out a remote Git branch?
- How do I force "git pull" to overwrite local files?
- How do I list all files of a directory?
- How to check whether a string contains a substring in JavaScript?
- How do I redirect to another webpage?
- How can I iterate over rows in a Pandas DataFrame?
- How do I convert a String to an int in Java?
- Does Python have a string 'contains' substring method?
- How do I check if a string contains a specific word?
Popular Tags
Trending Questions
- UIImageView Frame Doesn't Reflect Constraints
- Is it possible to use adb commands to click on a view by finding its ID?
- How to create a new web character symbol recognizable by html/javascript?
- Why isn't my CSS3 animation smooth in Google Chrome (but very smooth on other browsers)?
- Heap Gives Page Fault
- Connect ffmpeg to Visual Studio 2008
- Both Object- and ValueAnimator jumps when Duration is set above API LvL 24
- How to avoid default initialization of objects in std::vector?
- second argument of the command line arguments in a format other than char** argv or char* argv[]
- How to improve efficiency of algorithm which generates next lexicographic permutation?
- Navigating to the another actvity app getting crash in android
- How to read the particular message format in android and store in sqlite database?
- Resetting inventory status after order is cancelled
- Efficiently compute powers of X in SSE/AVX
- Insert into an external database using ajax and php : POST 500 (Internal Server Error)
Increase the dimensionality as desired, using 0 as new values.
From a theoretical point of view, consider the vector space as infinite dimensional.