A simple and efficient way to get document frequency counts of words from a corpus is to use CountVectorizer from Scikit Learn
Getting back to the word from the index is not immediately obvious, here's how to do it:
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer
docs = <load your docs as an iterable>
count_vect = CountVectorizer()
doc_counts = count_vect.fit_transform(docs) # this is of type scipy.sparse.csr.csr_matrix which is why we need to use
.ravel() below.
word_counts = zip(count_vect.get_feature_names(), np.asarray(doc_counts.sum(axis=0)).ravel())
word_counts = sorted(word_counts, key=lambda idx: -1 * idx[1] )
# Display top 100 words by frequency
word_counts[:100]
Tuesday, June 21, 2016
Thursday, June 9, 2016
How to 7zip each file seperately in a directory (Windows)
Let's zip up all those log files sitting in that directory into seperate 7z files....
FOR %i IN (*.*) DO "C:\Program Files\7-Zip\7z.exe" a -mx=9 "%i.7z" "%i"
Refs:
http://superuser.com/questions/312652/how-do-i-create-seperate-7z-files-from-each-selected-directory-with-7zip-command
http://askubuntu.com/questions/491223/7z-ultra-settings-for-zip-format
FOR %i IN (*.*) DO "C:\Program Files\7-Zip\7z.exe" a -mx=9 "%i.7z" "%i"
Refs:
http://superuser.com/questions/312652/how-do-i-create-seperate-7z-files-from-each-selected-directory-with-7zip-command
http://askubuntu.com/questions/491223/7z-ultra-settings-for-zip-format
Monday, April 18, 2016
Managing Java versions and Eclipse
Some little bits on managing Java versions and Eclipse:
Managing JRE installations in Eclipse
http://www.codejava.net/ides/eclipse/managing-jre-installations-in-eclipse
How to switch JDK version on Mac OS X
https://www.jayway.com/2014/01/15/how-to-switch-jdk-version-on-mac-os-x-maverick/
Edit your ~/.bash_profile and add the following:
function setjdk() {
if [ $# -ne 0 ]; then
removeFromPath '/System/Library/Frameworks/JavaVM.framework/Home/bin'
if [ -n "${JAVA_HOME+x}" ]; then
removeFromPath $JAVA_HOME
fi
export JAVA_HOME=`/usr/libexec/java_home -v $@`
export PATH=$JAVA_HOME/bin:$PATH
fi
}
function removeFromPath() {
export PATH=$(echo $PATH | sed -E -e "s;:$1;;" -e "s;$1:?;;")
}
setjdk 1.7
Managing JRE installations in Eclipse
http://www.codejava.net/ides/eclipse/managing-jre-installations-in-eclipse
How to switch JDK version on Mac OS X
https://www.jayway.com/2014/01/15/how-to-switch-jdk-version-on-mac-os-x-maverick/
Edit your ~/.bash_profile and add the following:
function setjdk() {
if [ $# -ne 0 ]; then
removeFromPath '/System/Library/Frameworks/JavaVM.framework/Home/bin'
if [ -n "${JAVA_HOME+x}" ]; then
removeFromPath $JAVA_HOME
fi
export JAVA_HOME=`/usr/libexec/java_home -v $@`
export PATH=$JAVA_HOME/bin:$PATH
fi
}
function removeFromPath() {
export PATH=$(echo $PATH | sed -E -e "s;:$1;;" -e "s;$1:?;;")
}
setjdk 1.7
Monday, March 21, 2016
Installing Java 8 on Ubuntu
Quick, easy point of reference for installing Java 8 on Ubuntu:
http://stackoverflow.com/questions/25549492/install-jdk8-in-ubuntu-14-04
sudo add-apt-repository ppa:webupd8team/java
sudo apt-get update
sudo apt-get install oracle-java8-installer
http://askubuntu.com/questions/315646/update-java-alternatives-vs-update-alternatives-config-java
sudo update-alternatives --config java
http://stackoverflow.com/questions/25549492/install-jdk8-in-ubuntu-14-04
sudo add-apt-repository ppa:webupd8team/java
sudo apt-get update
sudo apt-get install oracle-java8-installer
http://askubuntu.com/questions/315646/update-java-alternatives-vs-update-alternatives-config-java
sudo update-alternatives --config java
Tuesday, February 23, 2016
Debugging spark job incorrect JVM jar file loaded
I recently had an issue with incorrect version of a JVM jar file being loaded in a Spark job on Hadoop.
This Scala code snippet helped debug the issue and determine the loaded jar file path:
val jarPath = classOf[MyObject].getProtectionDomain().getCodeSource().getLocation().getPath()
In my case this pointed to an old version of guava:
/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-11.0.jar
Then by setting spark.driver.extraClassPath and spark.executor.extraClassPath arguments for spark-submit, the correct version of the jar file was loaded successfully:
spark-submit --class com.MyClass <other_spark_args> --conf "spark.driver.extraClassPath=/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-15.0.jar" --conf "spark.executor.extraClassPath=/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-15.0.jar" /home/mypath/myjarfile.jar <my_job_params>
For more info with extraClassPath, see: http://spark.apache.org/docs/latest/configuration.html
This Scala code snippet helped debug the issue and determine the loaded jar file path:
val jarPath = classOf[MyObject].getProtectionDomain().getCodeSource().getLocation().getPath()
In my case this pointed to an old version of guava:
/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-11.0.jar
Then by setting spark.driver.extraClassPath and spark.executor.extraClassPath arguments for spark-submit, the correct version of the jar file was loaded successfully:
spark-submit --class com.MyClass <other_spark_args> --conf "spark.driver.extraClassPath=/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-15.0.jar" --conf "spark.executor.extraClassPath=/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-15.0.jar" /home/mypath/myjarfile.jar <my_job_params>
For more info with extraClassPath, see: http://spark.apache.org/docs/latest/configuration.html
Sunday, November 1, 2015
Convert binary word2vec model to text vectors
If you have a binary model generated from google's awesome and super fast word2vec word embeddings tool, you can easily use python with gensim to convert this to a text representation of the word vectors.
Input: binary word embedding model from google's word2vec tool
Output: text vectors for word embeddings
Python conversion code:
from gensim.models import word2vec
model = word2vec.Word2Vec.load_word2vec_format('path/to/mymodel.bin', binary=True)
model.save_word2vec_format('path/to/mymodel.txt', binary=False)
I recommend using Anaconda from Continuum Analytics for a bundled python distribution. To install gensim in Anaconda just type: conda install gensim :)
Original ref: https://www.kaggle.com/c/word2vec-nlp-tutorial/forums/t/13828/how-to-convert-bin-file-of-word2vec-model-into-txt-r/91564
Input: binary word embedding model from google's word2vec tool
Output: text vectors for word embeddings
Python conversion code:
from gensim.models import word2vec
model = word2vec.Word2Vec.load_word2vec_format('path/to/mymodel.bin', binary=True)
model.save_word2vec_format('path/to/mymodel.txt', binary=False)
I recommend using Anaconda from Continuum Analytics for a bundled python distribution. To install gensim in Anaconda just type: conda install gensim :)
Original ref: https://www.kaggle.com/c/word2vec-nlp-tutorial/forums/t/13828/how-to-convert-bin-file-of-word2vec-model-into-txt-r/91564
Thursday, October 22, 2015
Analyse consecutive timeseries pairs in Spark
With time series data it is useful to construct pairs of elements for analysis.
Here is a way to construct a Spark RDD from a time series that has the pairs together in the final RDD:
val arr = Array((1, "A"), (8, "D"), (7, "C"), (3, "B"), (9, "E"))
val rdd = sc.parallelize(arr)
val sorted = rdd.sortByKey(true)
val zipped = sorted.zipWithIndex.map(x => (x._2, x._1))
val pairs = zipped.join(zipped.map(x => (x._1 - 1, x._2))).sortBy(_._1)
Which produces the consecutive elements as pairs in the RDD for further processing:
(0,((1,A),(3,B)))
(1,((3,B),(7,C)))
(2,((7,C),(8,D)))
(3,((8,D),(9,E)))
Ref: https://www.mail-archive.com/user@spark.apache.org/msg39353.html
Here is a way to construct a Spark RDD from a time series that has the pairs together in the final RDD:
val arr = Array((1, "A"), (8, "D"), (7, "C"), (3, "B"), (9, "E"))
val rdd = sc.parallelize(arr)
val sorted = rdd.sortByKey(true)
val zipped = sorted.zipWithIndex.map(x => (x._2, x._1))
val pairs = zipped.join(zipped.map(x => (x._1 - 1, x._2))).sortBy(_._1)
Which produces the consecutive elements as pairs in the RDD for further processing:
(0,((1,A),(3,B)))
(1,((3,B),(7,C)))
(2,((7,C),(8,D)))
(3,((8,D),(9,E)))
Ref: https://www.mail-archive.com/user@spark.apache.org/msg39353.html
Subscribe to:
Posts (Atom)