What data set is being used? Where does it come from, and what are the characteristics of it
(size, missing values, and continuous v/s categorical)?
I have decided on using Text Mining dataset. I would focus on sizing and continuous format.
Is there a reason you picked this data set? Tell me
Organizations utilize information and content mining to examine client and contender
information to enhance intensity. Advantages include: productivity; opening shrouded data and
growing new learning; investigating new skylines; enhanced research and confirmation base; and
enhancing the examination procedure and quality
More extensive financial and societal advantages were additionally featured, for example, cost
reserve funds and profitability increases, creative new administration improvement, new plans of
action and new restorative medicines.
Where content mining investigates copyrighted materials, the copyright holders may require
additional installment to enable their material to be utilized as a part of content mining. This is
notwithstanding the buy of the privilege to see the materials. In reality at times the client (or more
probable their organization) may need to pay four unique expenses to empower the materials to
be content mined conventional access (perusing) costs, the privilege to duplicate, the privilege
to digitize and afterward the privilege to content mine. As a few of those counseled featured, this
implies most content mining is restricted to investigating Open Access archives where no extra
charges are brought about.
Exchange costs in this setting identify with the exertion required to empower content mining to
occur. This is chiefly connected with getting authorization to mine specific corpora of reports. As