Popularity

7.5

Stable

Activity

0.0

Stable

Stars 553

Watchers 70

Forks 147

Last Commit over 6 years ago

Programming language: Scala

License: Apache License 2.0

Tags: Science And Data Analysis

Latest version: v1.2

FACTORIE alternatives and similar packages

Based on the "Science and Data Analysis" category.
Alternatively, view FACTORIE alternatives based on common mentions on social networks and blogs.

MLLib

10.0 10.0 FACTORIE VS MLLib

Apache Spark - A unified analytics engine for large-scale data processing
PredictionIO

9.9 0.0 FACTORIE VS PredictionIO

PredictionIO, a machine learning server for developers and ML engineers.

WorkOS - The modern identity platform for B2B SaaS

The APIs are flexible and easy-to-use, supporting authentication, user identity, and complex enterprise features like SSO and SCIM provisioning.

Promo workos.com

Zeppelin

9.8 8.7 L2 FACTORIE VS Zeppelin

Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.
Smile

9.7 9.8 L2 FACTORIE VS Smile

Statistical Machine Intelligence & Learning Engine
BigDL

9.7 9.9 FACTORIE VS BigDL

Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, Baichuan, Mixtral, Gemma, etc.) on Intel CPU and GPU (e.g., local PC with iGPU, discrete GPU such as Arc, Flex and Max). A PyTorch LLM library that seamlessly integrates with llama.cpp, HuggingFace, LangChain, LlamaIndex, DeepSpeed, vLLM, FastChat, ModelScope, etc.
Breeze

9.5 5.1 FACTORIE VS Breeze

Breeze is a numerical processing library for Scala.
Spark Notebook

9.5 0.0 L1 FACTORIE VS Spark Notebook

Interactive and Reactive Data Science using Scala and Spark.
Algebird

9.3 7.6 FACTORIE VS Algebird

Abstract Algebra for Scala
Spire

8.9 6.0 FACTORIE VS Spire

Powerful new number types and numeric abstractions for Scala.
Tensorflow_scala

8.0 0.0 FACTORIE VS Tensorflow_scala

TensorFlow API for the Scala Programming Language
Figaro

8.0 0.0 FACTORIE VS Figaro

Figaro Programming Language and Core Libraries
Squants

7.8 3.2 FACTORIE VS Squants

The Scala API for Quantities, Units of Measure and Dimensional Analysis
Saddle

6.9 0.0 FACTORIE VS Saddle

A minimalist port of Pandas to Scala
ND4S

6.1 0.0 FACTORIE VS ND4S

ND4S: N-Dimensional Arrays for Scala. Scientific Computing a la Numpy. Based on ND4J.
Chalk

5.7 0.0 FACTORIE VS Chalk

Chalk is a natural language processing library.
Compute.scala

4.8 0.0 FACTORIE VS Compute.scala

Scientific computing with N-dimensional arrays
Libra

4.5 0.0 FACTORIE VS Libra

A dimensional analysis library based on dependent types
Numsca

4.4 2.7 FACTORIE VS Numsca

numsca is numpy for scala
OpenMOLE

4.4 9.4 FACTORIE VS OpenMOLE

Workflow engine for exploration of simulation models using high throughput computing
Clustering4Ever

4.0 0.0 FACTORIE VS Clustering4Ever

C4E, a JVM friendly library written in Scala for both local and distributed (Spark) Clustering.
Optimus * 96

4.0 0.0 FACTORIE VS Optimus * 96

Optimus is a mathematical programming library for Scala.
rscala

3.6 6.1 FACTORIE VS rscala

The Scala interpreter is embedded in R and callbacks to R from the embedded interpreter are supported. Conversely, the R interpreter is embedded in Scala.
LoMRF

3.2 0.0 FACTORIE VS LoMRF

LoMRF is an open-source implementation of Markov Logic Networks
Tyche

2.9 0.0 FACTORIE VS Tyche

Statistics utilities for the JVM - in Scala!
MGO

2.8 5.6 FACTORIE VS MGO

Purely functional genetic algorithms for multi-objective optimisation
Rings

2.7 3.5 FACTORIE VS Rings

Rings: efficient JVM library for polynomial rings
Synapses

2.5 0.0 FACTORIE VS Synapses

A group of neural-network libraries for functional and mainstream languages
Axle

2.4 5.5 FACTORIE VS Axle

Axle Domain Specific Language for Scientific Cloud Computing and Visualization
SwiftLearner

2.2 0.0 FACTORIE VS SwiftLearner

SwiftLearner: Scala machine learning library
Persist-Units

0.9 0.0 FACTORIE VS Persist-Units

Scala Units of Measure Types
OscaR

0.2 - FACTORIE VS OscaR

a Scala toolkit for solving Operations Research problems

* Code Quality Rankings and insights are calculated and provided by Lumnify.
They vary from L1 to L5 with "L5" being the highest.

Do you think we are missing an alternative of FACTORIE or a related project?

Add another 'Science and Data Analysis' Package

Popular Comparisons

README

FACTORIE

This directory contains the source of FACTORIE, a toolkit for probabilistic modeling based on imperatively-defined factor graphs. More information, see the FACTORIE webpage.

Installation

Installation relies on Maven, version 3. If you don't already have maven, install it from http://maven.apache.org/download.html. Alternatively, you can use sbt as outlined below (a script for running sbt comes bundled with Factorie).

To compile type

$ mvn compile

To accomplish the same with sbt, type

$ ./sbt compile

You might need additional memory. If so, for sbt type

export SBT_OPTS="$SBT_OPTS -Xmx1g"

and for Maven type:

export MAVEN_OPTS="$MAVEN_OPTS -Xmx1g -XX:MaxPermSize=128m"

To create a self-contained .jar, that contains FACTORIE plus all its dependencies, including the Scala runtime, type

$ mvn -Dmaven.test.skip=true package -Pjar-with-dependencies

To accomplish the same with sbt, type

$ ./sbt assembly

To create a similar self-contained .jar that also contains all resources needed for NLP (including our lexicons and pre-trained model parameters), type

$ mvn -Dmaven.test.skip=true package -Pnlp-jar-with-dependencies

To accomplish the same with sbt, type

$ ./sbt -J-Xmx2G with-nlp-resources:assembly

Try out a simple example

To get an idea what a simple FACTORIE program might look like, open one of the class files in the tutorial package

$ ls src/main/scala/cc/factorie/tutorial

To run one of these examples using maven type

$ mvn scala:run -DmainClass=cc.factorie.tutorial.Grid

Try out implemented NLP models

Then you can run some FACTORIE tools from the command-line. For example, you can run many natural language processing tools.

$ bin/fac nlp --wsj-forward-pos --conll-chain-ner

will launch an NLP server that will perform part-of-speech tagging and named entity recognition in its input. The server listens for text on a socket, and spawns a parallel document processor on each request. To feed it input, type in a separate shell

$ echo "I told Mr. Smith to take a job at IBM in Raleigh." | nc localhost 3228

You can also run a latent Dirichlet allocation (LDA) topic model. Assume that "mytextdir" is a directory name containing many plain text documents each in its own file. Then typing

$ bin/fac lda --read-dirs mytextdir --num-topics 20 --num-iterations 100

will run 100 iterations of a sparse collapsed Gibbs sampling on all the documents, and print out the results every 10 iterations. FACTORIE's LDA implementation is faster than MALLET's.

You can also train a document classifier. Assume that "sportsdir" and "politicsdir" are each directories that contain plan text files in the categories sports and politics. Typing

$ bin/fac classify --read-text-dirs sportsdir,politicsdir --write-classifier mymodel.factorie

will train a log-linear by maximum likelihood (MaxEnt) and save it in the file "mymodel.factorie".

The above are simply a few simple command-line options. Internally the FACTORIE library contains extensive and general facilities for factor graphs: data representation, model structure, inference, learning.